Dev48
Language
  • About
  • Services
  • Industries
  • Technologies
  • Articles
  • Contacts
Book a call
    Home/Articles/Minimax m31 flash preview vs minimax m3 when the newer model is the riskier one
Dev48

© 2026 · All rights reserved.

MiniMax M3.1 Flash Preview vs MiniMax M3: When the Newer Model Is the Riskier One

Источник: OrcaRouter

MiniMax M3.1 Flash Preview vs MiniMax M3: When the Newer Model Is the Riskier One

Source: OrcaRouter

MiniMax-M3.1-Flash-Preview vs MiniMax-M3: the newer model adds a thinking-depth ladder but drops thinking-off, open weights and its published rate card.

September 28, 2026•Updated: September 28, 2026

The vendor's own documentation now lists MiniMax-M3.1-Flash-Preview above MiniMax-M3, and that ordering is doing a lot of persuasive work. It is the newest text model in the line, it carries the same million-token context window, and it adds a thinking-depth control the older model does not have. What the table does not show is that the older model is the only one of the two you can download, price per token, benchmark against anything, or call without a subscription — and that the newer one has taken away a switch that MiniMax-M3 gave you.

This is not a story about a model being superseded. It is a story about a vendor splitting its own line into a metered workhorse and a subscription-gated preview, and the two are not interchangeable in either direction. Here is what each of them actually is.

The two spec sheets, side by side

Start with what the vendor's own pages state for both, because the shape of the difference shows up here rather than in a benchmark table.

• Announced — MiniMax-M3 on 1 June 2026 with an architecture note, a benchmark sheet and a rate card; MiniMax-M3.1-Flash-Preview on 27 September 2026 with a documentation entry and a post on MiniMax's agent account • Context window — 1,000,000 tokens for both, per MiniMax's own model tables • Maximum output — 512,000 tokens for MiniMax-M3 on our catalogue entry, a figure MiniMax does not itself publish; not published for MiniMax-M3.1-Flash-Preview anywhere • Input modalities — text, image and video in, text out, for both • Thinking control — a thinking parameter MiniMax-M3 can be told to skip, and which can be separated into a reasoning_details field via reasoning_split; on MiniMax-M3.1-Flash-Preview thinking is mandatory and reasoning_split cannot be turned off • Thinking depth — not available on MiniMax-M3; an effort ladder of low, medium, high, xhigh and max, defaulting to max, on MiniMax-M3.1-Flash-Preview • Published output speed — roughly 100+ tokens per second for MiniMax-M3, per MiniMax's own model table; nothing published for MiniMax-M3.1-Flash-Preview • Weights — MiniMax-M3 is open-weight with a public repository; MiniMax-M3.1-Flash-Preview has none • Pricing — $0.30 per million input tokens and $1.20 per million output for MiniMax-M3 at the provider rate; no per-token rate exists for MiniMax-M3.1-Flash-Preview • Access — MiniMax-M3 on pay-as-you-go and on Token Plan, and routable through third-party platforms; MiniMax-M3.1-Flash-Preview on Token Plan and MiniMax Code only

Two rows there belong in a different category from the rest. The published output speed is a vendor number for the older model only, so the newer model is being sold as the fast one without a figure attached. And the thinking control has moved in the opposite direction from the marketing: the newer model offers finer control over how much it thinks while removing your ability to stop it thinking at all.

The switch MiniMax took away

This is the detail most likely to bite, and it is documented rather than hidden. With MiniMax-M3, sending thinking: {"type": "disabled"} on the OpenAI-compatible endpoint skips thinking and returns a direct answer; on the Anthropic-compatible endpoint, thinking is off by default and can be enabled with adaptive. On MiniMax-M3.1-Flash-Preview, the identical request returns HTTP 400 with the message requires adaptive thinking. There is no supported way to get a non-reasoning response.

Practically, that means a workload that used MiniMax-M3 as a cheap, fast classifier or router — short prompt in, short structured answer out, no chain of thought — has no drop-in upgrade path to the newer model. You can lower effort to low, and MiniMax's guidance is explicit that lowering effort is the intended way to reduce thinking latency and tokens rather than disabling it. But the tokens and the latency do not go to zero, and for a high-volume extraction job that writes four fields per call, "always thinks first" is a different cost shape from "thinks when asked".

What is actually measured on each

One side of this comparison has numbers and the other has a blank column, and it is worth being precise about which numbers are whose.

MiniMax-M3's independent record is real: Artificial Analysis puts it at 29.2 on the Intelligence Index — 60th of 145 — with 58.6 on the Coding Index, 59.18 on its Math index, 92.9% on GPQA Diamond under the same index family, 83% long-context recall and 82.9% on IFBench. Those are third-party evaluations of a model anybody can download and re-run. MiniMax's own launch numbers for M3 — BrowseComp at 83.5%, SWE-Bench Pro at 59.0%, Terminal-Bench 2.1 at 66.0%, MCP Atlas at 74.2%, OSWorld-Verified at 70% — are vendor figures and we have not seen them reproduced.

MiniMax-M3.1-Flash-Preview has neither. MiniMax published no benchmark table, and the model has no entry on Artificial Analysis's index, so there is no independent score, no coding index, no long-context recall figure and no hallucination measurement. Every characterisation of it in circulation — "fast", "reliable", "built for everyday development" — is the vendor's own adjective from the launch post. The absence is worth stating plainly because it cuts both ways: nothing has been shown to be worse, and nothing has been shown to be better.

That asymmetry is the whole decision. You can evaluate MiniMax-M3 this afternoon by downloading it or by calling it, and you can compare that evaluation to a public scoreboard. With MiniMax-M3.1-Flash-Preview you can only use it and form a private opinion.

Where the newer model genuinely wins

It would be easy to read the above as an argument for ignoring the new model, and that is not the right conclusion either. Three things are real and specific to it.

The effort ladder is a genuine capability, not a toggle. Five levels — low through max — let one model serve a routine rename and a repository-wide refactor at different compute costs, and it is documented as effective for MiniMax-M3.1-Flash-Preview alone. On MiniMax-M3 your choice is binary. If your problem is that you cannot predict per-task latency, a tunable depth is exactly the control you wanted.

It is the vendor's current top-of-table model, which means it inherits whatever the M-series lineage has learned since June, and MiniMax is explicit that it is the latest model for agentic reasoning, tool use, coding and long-context tasks. The interleaved-thinking tool-use machinery M3 introduced carries over, including the requirement to append the full response — thinking blocks and all — back into the message history.

And it takes video, at a Flash price point, inside a subscription. MiniMax-M3 does too, at a per-token price. If you are already paying for a Token Plan seat and your only barrier to using video input was marginal cost per call, M3.1-Flash-Preview removes that barrier.

Where the older model is still the one to call

Three cases, all of them about predictability rather than quality.

When you need to model cost. MiniMax-M3 has a rate card: $0.30 per million input tokens, $1.20 per million output, and a cache-read rate of $0.06 per million on our catalogue. You can estimate a workload's monthly bill to the cent before you run it. MiniMax-M3.1-Flash-Preview has no per-token rate anywhere on MiniMax's pages, so its cost is your subscription divided by however much you use it — fine for a fixed budget, useless for unit economics.

When you need the model to be the model. MiniMax-M3 is open-weight: you can download it, run it on your own hardware, and know that the artifact answering your calls in six months is the same one answering them today. MiniMax-M3.1-Flash-Preview has no checkpoint, and preview-grade access means the serving stack, quota composition and availability are all vendor-controlled and explicitly subject to change.

When you need a non-reasoning answer. As above — this is the only capability regression in the comparison, and it is the one that surprises people, because a newer model that cannot do something the older one could is not a case anyone plans for.

Calling both from one place

The awkwardness here is that the two models are bought differently — one metered, one by subscription — and the natural instinct is to pick a side. The more useful move is to route between them by task shape, which only works if the metered side is reachable from the same place as everything else.

MiniMax-M3 on OrcaRouter carries MiniMax's list price with 0% markup added, so the $0.30 and $1.20 on the rate card are what a call costs, with no platform margin on top. It sits on the same OpenAI-compatible endpoint as MiniMax M2.7 — the same $0.30/$1.20 sticker over a 200,000-token window, also open-weight — and as 200+ other models, behind one API key, with automatic failover across providers when one endpoint wobbles. Against our own telemetry, which is a different kind of number from a vendor's: over the last seven days M3 has answered at roughly 238 output tokens per second at a p50 of about 4.1 seconds to first token, with an error rate near 0.15%, measured on the OrcaRouter playground. If you want to hold MiniMax-M3 as the dependable half of a pipeline and treat MiniMax-M3.1-Flash-Preview as the experimental half, the dependable half is one line of configuration rather than a second vendor relationship. MiniMax-M3.1-Flash-Preview itself is not on our catalogue; the route to it is the vendor's own Token Plan and coding product.

What would change this article is specific and easy to watch for. A published rate card for MiniMax-M3.1-Flash-Preview would collapse the main reason to prefer M3 on economics. A set of independent scores would replace the blank column with something comparable. A downloadable checkpoint would remove the reproducibility objection. Until at least one of those arrives, MiniMax-M3 remains the model in this pair you can hold to account — and the newer one remains the faster, better-controlled, unmeasured preview that MiniMax has decided to charge for by the month instead of the token.

← All articles

More in AI & Machine Learning

All →
Lambda to build new data center in Mayes County, Oklahoma, generating half a billion dollars in tax revenue over next decade
Lambda

Lambda to build new data center in Mayes County, Oklahoma, generating half a billion dollars in tax revenue over next decade

OpenAI expands review of model behavior after more rogue agent incidents emergeПресса
OpenAI

OpenAI expands review of model behavior after more rogue agent incidents emerge

Apple faces $5.7 billion patent infringement verdict over iPhone and Apple Watch haptics
Пресса
Apple

Apple faces $5.7 billion patent infringement verdict over iPhone and Apple Watch haptics

Unsecured OpenAI agents posted 53 user images on the internet without the lab’s knowledgeПресса
OpenAI

Unsecured OpenAI agents posted 53 user images on the internet without the lab’s knowledge

Proaction boosts sales 60% and saves 75+ hours with Codex
OpenAI

Proaction boosts sales 60% and saves 75+ hours with Codex

Amazon data center communities: Here’s what’s happening near data centers across the US
Amazon

Amazon data center communities: Here’s what’s happening near data centers across the US

More from OrcaRouter

Claude Opus 5.5 Animated a Kākāpō Party in One HTML File: What the Demo Actually Ships
OrcaRouter

Claude Opus 5.5 Animated a Kākāpō Party in One HTML File: What the Demo Actually Ships

OpenAI's "GPT Gov" Claim, Checked: What's Actually Running in Classified Networks
OrcaRouter

OpenAI's "GPT Gov" Claim, Checked: What's Actually Running in Classified Networks

MiniMax M3.1 Flash Preview Is Live on the Token Plan: What Your Subscription Actually Unlocks
OrcaRouter

MiniMax M3.1 Flash Preview Is Live on the Token Plan: What Your Subscription Actually Unlocks

MiniMax M3.1 Flash Preview vs DeepSeek V4 Flash: Two Cheap 1M-Context Tickets, One Huge Asterisk
OrcaRouter

MiniMax M3.1 Flash Preview vs DeepSeek V4 Flash: Two Cheap 1M-Context Tickets, One Huge Asterisk