Dev48
Language
  • About
  • Services
  • Industries
  • Technologies
  • Articles
  • Contacts
Book a call
    Home/Articles/Minimax m31 flash preview is live on the token plan what your subscription actua
Dev48

© 2026 · All rights reserved.

MiniMax M3.1 Flash Preview Is Live on the Token Plan: What Your Subscription Actually Unlocks

Источник: OrcaRouter

MiniMax M3.1 Flash Preview Is Live on the Token Plan: What Your Subscription Actually Unlocks

Source: OrcaRouter

MiniMax-M3.1-Flash-Preview is live on the Token Plan: the same Subscription Key as M3, thinking always on, and no published price, weights or benchmarks.

September 28, 2026•Updated: September 28, 2026

MiniMax-M3.1-Flash-Preview went live on 27 September 2026, and the way it went live is the story. It is not a pay-as-you-go model. The vendor has put its newest text model behind the same Token Plan Subscription Key that already covers MiniMax-M3, MiniMax M2.7 and the M2.7-highspeed variant, so if you hold a seat there is no new key to cut and no second contract to sign. What there also is not, as of today, is a per-token price, a model card, a published benchmark, or any way to switch the model's thinking off.

That combination — frictionless access paired with entirely unpublished economics — is unusual enough to be worth unpacking before anyone wires a production pipeline to it. The rest of this piece is what the vendor's own documentation actually says, what it conspicuously does not say, and the one line of code that will tell you within a minute whether this model fits how you work.

What actually landed on 27 September

MiniMax's developer documentation now carries a standing note, repeated on every text-model page: MiniMax-M3.1-Flash-Preview is available only through Token Plan and MiniMax Code for now. The model appears first in the vendor's own model table, above MiniMax-M3, described as a frontier multimodal coding model with a 1M context window and tunable thinking depth.

• Context window — 1,000,000 tokens, the same ceiling MiniMax-M3 carries• Input modalities — text, image and video, returning text• Thinking depth — an effort setting accepting low, medium, high, xhigh and max, defaulting to max when omitted• Protocol support — the same model id works over MiniMax's Anthropic-compatible endpoint, its OpenAI-compatible endpoint and its OpenAI Responses endpoint• Availability — Token Plan and MiniMax Code only; not on pay-as-you-go• Price — no per-token rate published anywhere• Weights — none; the MiniMaxAI organisation's two most recent public repositories remain MiniMax-Music3 and MiniMax-H3• Independent benchmarks — none; the model has no entry on Artificial Analysis's model index

The company announced it the same day on its MiniMaxAgent account, framing it for everyday development rather than frontier work: fast and reliable, from quick bug fixes to full feature work. That is a Flash tier's brief, not a flagship's. Third-party wire coverage from PANews and KuCoin described it in the same terms and noted that developers had spotted the model in the MiniMax Code picker before the company confirmed it.

Which key you use matters more than usual

This is the part that catches people out. MiniMax runs two separate credential types, and M3.1-Flash-Preview only answers to one of them. The Token Plan Subscription Key comes from Billing > Token Plan and covers subscription usage and purchased Credits. The pay-as-you-go API Key covers the per-token models. The vendor documentation is explicit that the two are not interchangeable, and that a Subscription Key can exist before any paid resource is attached to it — in which case it is a valid key with nothing behind it.

So the practical test is not "do I have a MiniMax API key" but "do I have a Token Plan seat". Existing seats are covered. The documentation notes that a Subscription Key becomes usable when a seat or Credits access is assigned, and that Credits overflow is drawn automatically once the subscription quota is exhausted.

There is also a documentation inconsistency worth knowing about, because it will confuse anybody reading the pricing page first. The Token Plan pricing page's coverage footnote still lists the eligible lineup as M3, M2.7, image and speech, with H3 excluded — it has not been updated to name M3.1-Flash-Preview. The model pages say the opposite: that M3.1-Flash-Preview is available on Token Plan and nothing else. Both pages are live vendor pages as of today. Only the key in your hand settles it.

The 400 that tells you what kind of model this is

Every MiniMax M-series model thinks before it answers, but M3.1-Flash-Preview is the first one that refuses to stop. Send thinking: {"type": "disabled"} or reasoning_effort: "none" and the API returns HTTP 400 with the message:

model "MiniMax-M3.1-Flash-Preview" requires adaptive thinking; thinking.type="disabled" (including reasoning.effort=none) is not allowed (2013)

The vendor's own guidance is to lower effort rather than try to disable thinking, and the ladder is the latency lever: a routine edit at low and a repository-wide change at xhigh are the same model making different trades. Two further details are specific to this model. The reasoning_effort parameter is documented as effective for MiniMax-M3.1-Flash-Preview only — it is not a general M-series control. And the reasoning_split output switch, which separates thinking into a reasoning_content field, cannot currently be set to false on this model.

Where the thinking text lands is protocol-dependent: the Anthropic-compatible endpoint returns it as a thinking content block, the OpenAI-compatible endpoint returns it on reasoning_content, and the Responses endpoint returns it as a reasoning output item. If you were parsing <think> tags out of MiniMax-M3 output, that work is no longer necessary on the newer model — and for interleaved tool-use conversations, the full response including the thinking block has to be appended back into the message history or the reasoning chain breaks.

MiniMax Code is running a check-in promotion from 28 September to 7 October in UTC+8, doubling the free credits awarded for daily logins and opening it to new and existing users alike. Reported credits apply across supported models, with M3.1-Flash-Preview and the H3 video models named as examples. Separately, the company said Token Plan quotas reset for all users at launch and that further resets are planned during the event, with eligibility shown in-app.

Treat those two as vendor statements relayed through the launch coverage — they are promotional terms with dates attached, not product behaviour, and the same coverage notes that the announcement did not specify the per-check-in credit count or the precise rules for missed days. The steady-state economics are easier to state: Token Plan lists at $22 per month for Plus, $55 for Max and $132 for Ultra, with 5-hour rolling and weekly quota windows, sized by the vendor at roughly three to four, four to five, and six to seven concurrent agents respectively. Purchased Credits run at 1,000 credits to the dollar and deduct at each model's pay-as-you-go list price.

The four things this model still does not have

None of the following is a criticism of the model; each is a hole a team has to plan around, and all four are verifiable today.

• No price. There is no per-token rate for M3.1-Flash-Preview on any MiniMax page, so it cannot be compared with M3's $0.30 per million input and $1.20 per million output, or with anything else, on cost per token.• No benchmarks. MiniMax published no evaluation table with this release. Neither has anyone else: the model is absent from Artificial Analysis's index, so its column is genuinely empty rather than merely early.• No weights. MiniMax-M3 shipped open-weight; M3.1-Flash-Preview has no repository and no downloadable checkpoint.• No independent measurement. No output-speed, latency or error-rate figures from any third party, and none from our own routing telemetry, because the model is not on our catalogue.

Against the M3 launch in June 2026, which arrived with an architecture note, a benchmark sheet and a rate card on day one, this is a deliberately quiet release. The plausible read — and it is a read, not a stated plan — is that MiniMax wants real usage inside its own product before it decides what the model costs and whether it opens up.

What the leaked preview note got right

Four days before the launch, a preview document transcribed in a public partner repository circulated describing a MiniMax M3.1 with sparse attention, Q8KV4 attention quantisation, NVFP4 routed experts, a DSpark speculative-decoding head and a new reasoning_effort field taking values from max down to low. At the time we wrote it up as unverified, because it was: no weights, no model card, no pricing page and no API model id existed.

One of those five claims is now confirmed by the vendor's own documentation, and it is the reasoning_effort field — including the max-to-low ladder, which matches the shipped values exactly. The other four remain unconfirmed. MiniMax's published description of M3.1-Flash-Preview says only that it is a frontier multimodal coding model with a 1M context window and tunable thinking depth; it does not describe the attention scheme, the quantisation or the decoding head. Anyone repeating the sparse-attention and NVFP4 claims as settled specification is repeating a third party's transcription, which is exactly what the release did not confirm.

If you want to try it without betting a pipeline on it

The awkward part of a Token-Plan-only preview is that the natural way to evaluate a new model — call it from your existing stack and measure — is the one thing you cannot do cleanly. There is no per-token rate to model against, no published latency profile, and no failover story inside the vendor's own product if the preview wobbles.

That is the situation a routing layer is built for. Everything you can call today lives behind one OpenAI-compatible endpoint and one API key, and the models adjacent to this one are already there: MiniMax-M3 at the provider's rate of $0.30 per million input tokens and $1.20 per million output, MiniMax M2.7 at the same sticker with a 200,000-token window and open weights, and the DeepSeek and Qwen Flash tiers in the same price band. Prices on our side are the provider's list price passed through at 0% markup, so when a vendor cuts a rate it lands here the same day rather than on a renegotiation cycle. Automatic failover means a preview-grade dependency can sit beside a production path rather than under it, and when MiniMax does publish a rate card and an endpoint for MiniMax-M3.1-Flash-Preview, the question becomes a routing-table entry instead of a migration project. MiniMax-M3.1-Flash-Preview itself is not on our catalogue, and the way to reach it today is the vendor's own Token Plan and its coding product.

The honest summary for a team reading this on launch day: the model is live, free at the margin if you already pay for a Token Plan seat, and completely unpriced and unmeasured on its own terms. Take it for a spin inside MiniMax Code, keep the pipeline you depend on on a model with a rate card and a track record, and watch for the two things that will turn this from a curiosity into a decision — a published price and the first independent scores.

← All articles

More in AI & Machine Learning

All →
Lambda to build new data center in Mayes County, Oklahoma, generating half a billion dollars in tax revenue over next decade
Lambda

Lambda to build new data center in Mayes County, Oklahoma, generating half a billion dollars in tax revenue over next decade

OpenAI expands review of model behavior after more rogue agent incidents emergeПресса
OpenAI

OpenAI expands review of model behavior after more rogue agent incidents emerge

Apple faces $5.7 billion patent infringement verdict over iPhone and Apple Watch haptics
Пресса
Apple

Apple faces $5.7 billion patent infringement verdict over iPhone and Apple Watch haptics

Unsecured OpenAI agents posted 53 user images on the internet without the lab’s knowledgeПресса
OpenAI

Unsecured OpenAI agents posted 53 user images on the internet without the lab’s knowledge

Proaction boosts sales 60% and saves 75+ hours with Codex
OpenAI

Proaction boosts sales 60% and saves 75+ hours with Codex

Amazon data center communities: Here’s what’s happening near data centers across the US
Amazon

Amazon data center communities: Here’s what’s happening near data centers across the US

More from OrcaRouter

Claude Opus 5.5 Animated a Kākāpō Party in One HTML File: What the Demo Actually Ships
OrcaRouter

Claude Opus 5.5 Animated a Kākāpō Party in One HTML File: What the Demo Actually Ships

OpenAI's "GPT Gov" Claim, Checked: What's Actually Running in Classified Networks
OrcaRouter

OpenAI's "GPT Gov" Claim, Checked: What's Actually Running in Classified Networks

MiniMax M3.1 Flash Preview vs DeepSeek V4 Flash: Two Cheap 1M-Context Tickets, One Huge Asterisk
OrcaRouter

MiniMax M3.1 Flash Preview vs DeepSeek V4 Flash: Two Cheap 1M-Context Tickets, One Huge Asterisk

MiniMax M3.1 Flash Preview vs MiniMax M3: When the Newer Model Is the Riskier One
OrcaRouter

MiniMax M3.1 Flash Preview vs MiniMax M3: When the Newer Model Is the Riskier One