MAI-Transcribe-2-Streaming vs MAI-Transcribe-2: Pay 5.4× for Word-Level Timing

Источник: OrcaRouter

MAI-Transcribe-2-Streaming vs MAI-Transcribe-2: Pay 5.4× for Word-Level Timing

Source: OrcaRouter

Same Microsoft family, two tiers: streaming at $9.00 per 1,000 minutes vs batch at $1.67. Diarization and word timestamps ship only on the batch side.

•Updated: October 1, 2026

MAI-Transcribe-2-Streaming and MAI-Transcribe-2 are the same model family split across two pricing tiers, and the split is clean enough to plan against. MAI-Transcribe-2 shipped on September 3, 2026 as a batch model: 2.0% word error rate on Artificial Analysis's non-streaming board, roughly 411× real time, and $0.10 per audio-hour — about $1.67 per 1,000 minutes. MAI-Transcribe-2-Streaming shipped on October 1, 2026 as the real-time variant: a claimed 2.5% streaming word error rate at 0.13 seconds to final transcript, which Microsoft describes as first for accuracy on both final and partial transcripts, measured on Artificial Analysis's streaming methodology rather than yet plotted as a row on the tracker's board. $0.54 per audio-hour — about $9.00 per 1,000 minutes. Same 60 languages, same API-only distribution, no published weights. Streaming costs 5.4× the batch rate and, on Microsoft's own numbers, 0.5 points of accuracy. Whether that is a bargain or a tax depends entirely on one question: does anything in your pipeline need the transcript before the sentence is over?

Because the two models come from one vendor and one docs surface, the comparison is unusually clean — no feature-parity guessing, no benchmark-version mismatch, no "their numbers versus our numbers." It is a single vendor pricing two speeds of the same capability, and the interesting work is figuring out which workloads the cheap tier actually covers.

The family, in two rows

• MAI-Transcribe-2 — batch, pre-recorded audio, released September 3, 2026. 2.0% AA-WER, second on Artificial Analysis's non-streaming leaderboard. Roughly 411× real time, also second. $0.10 per audio-hour, about $1.67 per 1,000 minutes, as a limited-time offer through the end of the year with no standard price published. Bundles speaker diarization and returns word-level timestamps. 60 languages with automatic language identification.

• MAI-Transcribe-2-Streaming — real-time, WebSocket-style continuous transcription, released October 1, 2026. Microsoft reports 2.5% final-transcript WER at 0.13 seconds after end of speech, and the same 2.5% on the first partial at 0.12 seconds, which it describes as first for accuracy on both — a vendor claim measured on Artificial Analysis's streaming methodology, though as of drafting the tracker's own board (37 models) has not yet plotted the model, and its current leader is Grok Voice Transcribe 2.0 at 2.73% and 0.49 seconds. $0.54 per audio-hour, about $9.00 per 1,000 minutes, introductory through year-end. 60 languages with automatic continuous language detection. No published diarization claim, no published word-timestamp claim.

The one asymmetry that is not about speed or price is capability: the batch model's diarization and word timestamps are documented features, and the streaming model's are not. That matters more than the 0.5-point accuracy gap, and it is the reason a lot of teams will end up running both.

What 5.4× actually buys

At $1.67 per 1,000 minutes, batch transcription at this accuracy tier is close to free at the margin. At $9.00 per 1,000 minutes, streaming is a real line item — roughly the cost of the same audio transcribed five times over, plus a half-point of accuracy. So the question is not whether real-time is worth more; it is which specific minutes need it.

The workloads where it clearly is: live captions a person reads while the speaker talks, where 0.13 seconds versus end-of-file is the entire product; real-time agent assist and compliance scoring, where an intervention that arrives after the call ended is worthless; and interactive voice applications, where a turn cannot complete until the transcript does. In all three, the batch model is not a cheaper option, it is a non-option — there is no price at which post-hoc transcription becomes real-time.

The workloads where it clearly is not: meeting and call recording processed after the fact, media archives, contact-center batch QA, compliance review of completed calls, podcast and video captioning. These are exactly the pipelines where the batch model's 411× real time and $1.67 rate were built to win, and paying 5.4× to shave a few hundred milliseconds off a transcript nobody reads until tomorrow is a straightforward waste. Add the diarization and word-timestamp features — the batch model has them documented, the streaming model does not — and the post-hoc case gets stronger, not weaker.

The interesting middle is where the two are genuinely complementary: run streaming for the live surface and batch for the record. A support call transcribed in real time to drive an agent-assist panel, then re-transcribed in batch overnight to produce the archival transcript with diarization and timestamps, costs a small premium over batch alone and gets both behaviours. That pattern only became possible in September, and the October release is what completes it.

Where the two-division structure shows

Two things about this family are worth noticing because they are structural rather than incidental.

First, the accuracy hierarchy is inverted from the usual pattern. Streaming models normally pay a substantial accuracy tax for real-time output; here the tax is 0.5 points, from 2.0% on the batch board to a claimed 2.5% on the streaming one. Microsoft got a real-time model within half a point of its batch flagship on its own numbers, which is the actual engineering achievement of the October release if it holds — not the top-of-board placement, which a good re-run could move and which the tracker has not yet confirmed.

Second, the pricing is inverted from the usual pattern too. Vendors typically discount the higher-volume tier; Microsoft priced streaming at 5.4× batch and left both on introductory rates expiring at the same time. That reads less like a deliberate streaming premium and more like two separate launches each carrying a launch discount, with no published standard rate for either. Teams planning a 2027 budget should treat both numbers as provisional — the $1.67 and the $9.00 are both labeled limited-time, and neither has an announced successor price.

What is not proven yet

The batch model's open questions from September are still open: no published diarization error rate, no per-language breakdown behind the 60-language claim, and a promotional price with no successor named. The streaming model adds one more that is specific to real-time work — streaming WER is measured on clean audio, and partial-transcript quality in a noisy far-field stream is the failure mode that matters most in deployment and is least covered by the launch material. It also inherits the streaming board's closeness: the top three on the tracker's 37-model board sit within a fraction of a point, so "first" is a state that a re-run can change — and Microsoft's first is a claim rather than a row.

The family-level question is the one to watch, though. Microsoft now sells the same capability at two prices five times apart, with the expensive tier missing two features the cheap tier documents. If diarization and word timestamps arrive on the streaming SKU, the gap narrows to purely a latency-and-price trade, which is a much simpler decision. If they do not, the two-division structure is permanent and architectures should be built around it.

Where this leaves a transcription stack

The practical upshot is that transcription stopped being a single line item in September and became a routing decision in October. A pipeline that used to standardize on one API now wants streaming for the live surface, batch for the record, and a rule for which audio takes which path — and that rule is itself a routing problem, which is a strange amount of complexity to land in one vendor's release cycle.

It lands hardest on what sits above the transcript. Whichever tier produces the text, that text then gets summarized, classified, redacted, and routed, and those steps run on language models with their own latency and cost profiles — a live transcript wants a fast, cheap model, an archival one can afford a stronger one. OrcaRouter covers exactly that layer: one API across 200+ models, provider list prices passed through at 0% markup, automatic failover, and a routing DSL that picks a model per workload rather than per pipeline. Neither MAI-Transcribe-2-Streaming nor MAI-Transcribe-2 is on that roster — both are served through Microsoft's own speech stack — but a two-division transcription layer is the strongest argument yet for keeping everything downstream of it portable.

The honest summary: streaming is not a better MAI-Transcribe-2, it is a second product at a second price. Buy it for the minutes that genuinely need to be real-time, leave the rest on the cheap tier, and budget for both rates being re-priced when the introductory period ends.

What this article says

Something is unclear? Ask about the article — I will explain in plain words.

Do not want to dig deeper? We will sort it out for you.