Dev48
Language
  • About
  • Services
  • Industries
  • Technologies
  • Articles
  • Contacts
Book a call
    Home/Articles/Minimax m31 flash preview vs deepseek v4 flash two cheap 1m context tickets one
Dev48

© 2026 · All rights reserved.

MiniMax M3.1 Flash Preview vs DeepSeek V4 Flash: Two Cheap 1M-Context Tickets, One Huge Asterisk

Источник: OrcaRouter

MiniMax M3.1 Flash Preview vs DeepSeek V4 Flash: Two Cheap 1M-Context Tickets, One Huge Asterisk

Source: OrcaRouter

MiniMax-M3.1-Flash-Preview vs DeepSeek V4 Flash: both offer a 1M context. One has a readable rate card under a retired name; the other no published price.

September 28, 2026•Updated: September 28, 2026

If you are shopping for a million-token context window at a low price per token, MiniMax-M3.1-Flash-Preview and DeepSeek V4 Flash are the two names that keep coming up — and they are not really the same kind of product. DeepSeek V4 Flash is a retired model whose identifier still answers, living inside a vendor rate card you can read to the cent. MiniMax-M3.1-Flash-Preview is a live preview with no published rate at all. The comparison below is therefore partly a spec comparison and partly a demonstration of how differently these two vendors are choosing to sell the same bracket.

Both sides of it are worth stating precisely, because the numbers that decide it are scattered across a vendor pricing page, a vendor documentation tree and our own catalogue, and one of the most frequently repeated claims about DeepSeek V4 Flash — its price — is one DeepSeek has since replaced.

What the two spec sheets actually match on

The headline specs are close enough that they are not the deciding factor, with one exception nobody advertises.

• Context window — 1,000,000 tokens for MiniMax-M3.1-Flash-Preview against 1,048,576 for DeepSeek V4 Flash: nominally identical, a 48,576-token difference• Maximum output — not published for MiniMax-M3.1-Flash-Preview; 384,000 tokens for DeepSeek V4 Flash, which is unusual at this price and is the number that should decide a lot of purchasing• Input modalities — text, image and video for MiniMax-M3.1-Flash-Preview; text only for DeepSeek V4 Flash• Thinking control — an effort ladder from low to max, thinking permanently on, for MiniMax-M3.1-Flash-Preview; thinking or non-thinking on DeepSeek V4 Flash, with non-thinking as a real option• Protocol support — Anthropic-compatible, OpenAI-compatible and OpenAI Responses endpoints for MiniMax-M3.1-Flash-Preview; the same three plus chat prefix completion on DeepSeek V4 Flash• Weights — none for MiniMax-M3.1-Flash-Preview; DeepSeek's V4 Flash weights were published, though the model itself has been superseded• Price — unpublished for MiniMax-M3.1-Flash-Preview; published and low for DeepSeek V4 Flash

Read that list twice, because the output ceiling is the quiet one. A million tokens of context with 384,000 tokens of output is a genuinely long single-pass generation window. A million tokens of context with an undisclosed output ceiling is a promise about reading, not about writing. If your workload is "read a repository, emit a migration plan", the second number is the one that determines whether the task fits in one call.

What each one is measured at

This is the least comfortable section of the comparison, and it is short.

DeepSeek V4 Flash has a benchmark record, and it is independent. Artificial Analysis scores it 34.3 on the Intelligence Index — 42nd of 145 models on that index — with 69.1 on the Coding Index and 90.8% on GPQA Diamond. The 95.0 on τ²-Bench that our own catalogue entry leads with is the same figure DeepSeek's launch material carried, so it is a vendor number in independent company rather than an audited one. Its long-context recall is 79.7% and its terminal-bench figure is 78.7% on the v2.1 harness. Those are real, comparable measurements of a known model.

MiniMax-M3.1-Flash-Preview has none. MiniMax published no evaluation table with the release, and no third party has one either — the model does not appear on Artificial Analysis's index at all. That is not a gap in our research; it is the current state of the public record, and it means every quality claim about the newer model is, at this moment, a vendor adjective. Anyone who hands you a MiniMax-M3.1-Flash-Preview benchmark table today is either quoting the older MiniMax-M3 or making it up.

DeepSeek V4 Flash is not sold at the price you remember

Here is the trap. The figure widely quoted for DeepSeek V4 Flash — $0.14 per million input tokens and $0.28 per million output — was the card the model carried before 10 September 2026. On that date DeepSeek released DeepSeek-V4.1-Flash, retired both V4 Flash and the V4-Flash-Vision-Exp experiment, and cut prices. The deepseek-v4-flash identifier still works, but DeepSeek's documentation states that it is temporarily routed to V4.1 Flash and billed at the Flash price. In other words, you can still type the old model name; you are not calling the old model.

So the real card to compare against is the current one, per 1M tokens — and off-peak is the cheaper half of it:

• Cache-miss input — $0.15 off-peak, $0.30 at peak• Cache-hit input — $0.003 off-peak, $0.006 at peak• Output — $0.60 off-peak, $1.20 at peak• Peak windows — 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, excluding Chinese public holidays; everything else, including full weekends, is off-peak

Against that, MiniMax-M3.1-Flash-Preview has no per-token rate of any kind. Its cost is whatever your Token Plan seat costs — $22, $55 or $132 per month for Plus, Max and Ultra — burned down through a shared quota pool at each model's pay-as-you-go list price, with the vendor reserving the right to change the composition of that pool. For a fixed monthly budget with predictable volume that can be excellent value. For a project that needs to attribute cost per call, it is opacity dressed as a subscription.

The reason this matters more than usual on the DeepSeek side is the peak multiplier. A team in Europe or Asia running its working day is inside a peak window for a meaningful share of its traffic and pays double for the privilege, which turns a headline $0.15 input rate into $0.30 for exactly the hours most products are busy. Budget off the peak column unless your traffic genuinely runs at night in UTC.

Where MiniMax M3.1 Flash Preview is the wrong call

Three cases, and the first two are easy to miss because they are documented rather than subtracted.

The output ceiling is unknown. If your task ends in a long generation — a full specification, a transcript rewrite, an annotated report — you are choosing an output budget you cannot see. DeepSeek V4 Flash publishes 384,000. That asymmetry favours the older model for anything that writes a lot.

Thinking cannot be turned off. Send thinking: {"type": "disabled"} to MiniMax-M3.1-Flash-Preview and you get an HTTP 400: the model requires adaptive thinking. On DeepSeek V4 Flash, non-thinking mode exists and is a supported switch. For high-throughput classification, routing, or short structured extraction, a model that always reasons first is paying latency and tokens you did not ask for — and the only lever you have is dropping effort to low, not removing the reasoning.

It is not callable on pay-as-you-go. The vendor's documentation is unambiguous that MiniMax-M3.1-Flash-Preview is available only through Token Plan and MiniMax Code right now, and the Subscription Key is not interchangeable with a pay-as-you-go API key. If you already hold a seat, it is a one-line change. If you do not, evaluating this model means buying a subscription.

Where the DeepSeek number is the wrong call

The mirror image is that DeepSeek V4 Flash is text-only and already superseded. It never accepted images — that work needed the separate experimental build, which has now been retired alongside it — and the model identifier now serves different weights than the name suggests. A pipeline pinned to deepseek-v4-flash for reproducibility is relying on a redirect DeepSeek describes as temporary.

MiniMax-M3.1-Flash-Preview takes text, image and video natively, and it is the vendor's current top-of-table text model. Video input at a Flash price point is not common in this bracket.

Both caveats point the same way for anyone running production traffic: the identifier on the invoice and the model that answered are not guaranteed to be the same thing indefinitely, and a routing layer is where that gets handled rather than discovered.

How to decide this in one afternoon

Start from the ending of the task rather than the beginning.

If the job is long-form generation — anything that finishes by writing tens of thousands of tokens — DeepSeek V4 Flash is the only one of the two that publishes a ceiling big enough to promise it, and its rate card is small enough that the peak multiplier is survivable. Model the cost at peak, not off-peak, and treat the retired identifier as a migration you have not done yet rather than a stable contract.

If the job is reading and reasoning over multimodal input under a fixed monthly budget — screen recordings, scanned documents, video review — MiniMax-M3.1-Flash-Preview is the more capable shape, and inside a Token Plan seat it is the cheapest way to get video input at this context length. Accept that you are buying an unmeasured model, and cap the blast radius: keep the workload that matters on something with a published rate card, and let the preview handle the exploratory half.

If both descriptions fit your pipeline, that is a routing problem rather than a model-selection problem, and it is the case we exist for. DeepSeek V4 Flash on OrcaRouter sits at the provider's rate with 0% markup — its catalogue entry shows $0.22 per million input tokens and $0.66 per million output, which is the peak column of the current Flash card, passed straight through — alongside MiniMax-M3 and MiniMax M2.7 in the same price band on the same key. One endpoint reaches 200+ models with automatic failover behind it, so a workload can move between price brackets without a second integration. MiniMax-M3.1-Flash-Preview is not on our catalogue; reaching it means the vendor's Token Plan and its coding product.

The one-line version: DeepSeek V4 Flash is the known quantity with the published output ceiling, the readable rate card and a successor quietly answering to its name; MiniMax-M3.1-Flash-Preview is the newer, multimodal, unmeasured one that costs a subscription to try. Neither of those is a reason to pick the other — they are just the facts you were not going to get from a price-per-token comparison that quotes a card DeepSeek retired in September.

← All articles

More in AI & Machine Learning

All →
Lambda to build new data center in Mayes County, Oklahoma, generating half a billion dollars in tax revenue over next decade
Lambda

Lambda to build new data center in Mayes County, Oklahoma, generating half a billion dollars in tax revenue over next decade

OpenAI expands review of model behavior after more rogue agent incidents emergeПресса
OpenAI

OpenAI expands review of model behavior after more rogue agent incidents emerge

Apple faces $5.7 billion patent infringement verdict over iPhone and Apple Watch haptics
Пресса
Apple

Apple faces $5.7 billion patent infringement verdict over iPhone and Apple Watch haptics

Unsecured OpenAI agents posted 53 user images on the internet without the lab’s knowledgeПресса
OpenAI

Unsecured OpenAI agents posted 53 user images on the internet without the lab’s knowledge

Proaction boosts sales 60% and saves 75+ hours with Codex
OpenAI

Proaction boosts sales 60% and saves 75+ hours with Codex

Amazon data center communities: Here’s what’s happening near data centers across the US
Amazon

Amazon data center communities: Here’s what’s happening near data centers across the US

More from OrcaRouter

Claude Opus 5.5 Animated a Kākāpō Party in One HTML File: What the Demo Actually Ships
OrcaRouter

Claude Opus 5.5 Animated a Kākāpō Party in One HTML File: What the Demo Actually Ships

OpenAI's "GPT Gov" Claim, Checked: What's Actually Running in Classified Networks
OrcaRouter

OpenAI's "GPT Gov" Claim, Checked: What's Actually Running in Classified Networks

MiniMax M3.1 Flash Preview Is Live on the Token Plan: What Your Subscription Actually Unlocks
OrcaRouter

MiniMax M3.1 Flash Preview Is Live on the Token Plan: What Your Subscription Actually Unlocks

MiniMax M3.1 Flash Preview vs MiniMax M3: When the Newer Model Is the Riskier One
OrcaRouter

MiniMax M3.1 Flash Preview vs MiniMax M3: When the Newer Model Is the Riskier One