HeyGen Voice vs Qwen-Audio-3.0-TTS: What You Pay for the Elo, and Which Qwen Tier You Actually Benchmarked

Source: OrcaRouter•

HeyGen Voice vs Qwen-Audio-3.0-TTS: What You Pay for the Elo, and Which Qwen Tier You Actually Benchmarked

HeyGen Voice scores 1,202 at $30 per 1M characters; Qwen-Audio-3.0-TTS-Plus scores 1,123 at a published $19.30. The Qwen tier you pick matters.

Run the arithmetic first, because it is less lopsided than the scoreboard suggests. Qwen-Audio-3.0-TTS-Plus costs $19.30 per million characters on the Model Studio platform — a rate the vendor publishes. HeyGen Voice sits at $30.00 per million characters on Artificial Analysis's own price chart, a figure the site derives from HeyGen's credit pricing, because HeyGen's public pricing page sells credits without stating what a voice generation consumes. Meanwhile the controlled-voice gap between them is 79 Elo: HeyGen Voice scored 1,202 on the arena that holds all eight reference voices constant, Qwen-Audio-3.0-TTS-Plus 1,123. So the cheaper model is ranked well behind the better one, the ratio between them is roughly 1.55×, and only one of those two prices is a number the vendor selling it put its name to.

The trap in the name

Before any of that, get the tier right, because this is a family that will happily let you benchmark the wrong model. Two Alibaba entries matter here and they are not interchangeable.

Qwen-Audio-3.0-TTS-Plus is the model this article compares against. It scores 1,123 on the controlled board — seventh of forty-two — and 1,264 on the Provider Voices board, where it ranks sixth of ninety-six. It is a real, current, streaming TTS product with a published rate of $19.30 per million characters.

Qwen-Audio-3.1-TTS-Plus is its successor tier, and it is the one sitting second on the controlled board at 1,186 — the model that came closest to beating HeyGen Voice on debut. Its rate is listed at the same $19.30, though Alibaba's documentation warns that different model versions require matching voices, so "same price, newer tier" does not mean "drop-in".

Anyone who reads "#1 ahead of Qwen" and then benchmarks Qwen-Audio-3.0-TTS will be testing a model 63 Elo weaker than the one the headline was about. The tier quibble is worth 63 points, which is more than the gap between HeyGen Voice and third place.

The scoreboard

• Controlled Voice Elo (same 8 cloned voices) — HeyGen Voice: 1,202, #1 of 42. Qwen-Audio-3.0-TTS-Plus: 1,123, #7.

• Provider Voices Elo (native voices) — Qwen-Audio-3.0-TTS-Plus: 1,264, #6 of 96. HeyGen Voice: not ranked.

• Price per 1M characters — Qwen-Audio-3.0-TTS-Plus: $19.30, vendor-published on Alibaba Cloud. HeyGen Voice: $30.00 on the arena's own conversion of credit pricing; HeyGen publishes no voice rate.

• Cost of 100 hours of narration — Qwen-Audio-3.0-TTS-Plus: roughly $926 at ~480,000 characters per spoken hour. HeyGen Voice: roughly $1,440 on the arena's $30.00 conversion — derived, not quoted.

• Voice control — Qwen-Audio-3.0-TTS-Plus: instruction parameter, emotion and language tags, voice cloning and voice design. HeyGen Voice: text-to-voice design from a ≤1,000-character description, instant clone, professional clone trained on 20+ minutes.

• Output — Qwen-Audio-3.0-TTS-Plus: PCM, WAV, MP3 and Opus at up to 48kHz, with adjustable rate, pitch, volume and bitrate. HeyGen Voice: speed 0.5–1.5 and pitch ±50 semitones per request.

What Alibaba's engineering actually gives you

The Qwen-Audio-3.0-TTS family is a streaming product and the documentation reads like it. Requests go over a WebSocket protocol shared with CosyVoice, with the run-task, continue-task and finish-task event flow and task-started, result-generated, task-finished and task-failed responses alongside them. Cancellation is supported on every Qwen-Audio-TTS model in both the Beijing and Singapore regions via streaming_cancel(). Output reaches 48kHz, and the formats include Opus, which matters if you are moving audio into a browser or a phone call rather than a WAV file.

Two details in that documentation are easy to miss and expensive to discover late. Emotion and rich language tags apply only to the Plus and Flash tiers — not the whole family — and they work only in unidirectional streaming mode. And API keys differ between the Singapore and Beijing regions, with workspaces addressed at URLs of the form wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/inference. Regional keys are a provisioning detail until the day they are an outage.

HeyGen's side is a different shape of product: a REST-style v3 API where POST /v3/voices/speech renders speech, GET /v3/voices lists the catalogue behind an engine filter, and voice design returns up to three ranked options from a text prompt with a seed for reproducibility. The documentation also reveals that the engine filter accepts values beyond HeyGen's own — so HeyGen will happily route you to a third-party voice engine through its API. A vendor that ships a competitor's model as a selectable option is telling you something about where it thinks its own strength lies.

Where the 79 Elo goes

A 79-point gap on a controlled board is substantial — larger than the 16-point margin that separates HeyGen Voice from second place, and roughly the distance between third and eighth in that field. It is also measured on one axis.

The controlled arena clones eight voices and holds them fixed. It says nothing about the instruction-driven control surface Qwen exposes, nothing about 48kHz Opus output through a cancellable WebSocket, and nothing about regional deployment. Those are engineering properties, and they are the reason a team building a Mandarin-language contact centre would not switch to a model that scores 79 points higher on eight English-language reference voices.

What the controlled result does say is narrower and still useful: when both engines are handed the same cloned voice, listeners preferred HeyGen Voice's rendering. If your product clones voices, that is your axis. If your product picks a voice from a catalogue and streams it into a call, the provider board is closer to your reality — and there HeyGen Voice has no rank at all while Qwen-Audio-3.0-TTS-Plus sits sixth of ninety-six.

The price argument, taken seriously

At $19.30 per million characters, Qwen-Audio-3.0-TTS-Plus is roughly half the cost of ElevenLabs' $40 tier and about 40% of Cartesia Sonic 3.6's $49. For a workload of 100 hours of narration — around 48 million characters at a typical speaking rate — that is on the order of $926. The same volume on Eleven v4 Turbo would be about $1,920 and on Sonic 3.6 about $2,352. HeyGen Voice, at the arena's $30.00, lands near $1,440: between the cheap tier and the two incumbents, closer to the middle than the top.

That arithmetic carries a caveat the others do not need. $19.30 is Alibaba's own published rate. $30.00 is Artificial Analysis's conversion of a credit system HeyGen has never translated into characters, so it is a good-faith estimate rather than a quotation. If the real characters-per-credit ratio is worse than the chart assumes, the 79-Elo advantage costs more than the 1.55× it implies; if it is better, less. A model whose price you have to infer is a model you pilot on a measured slice of traffic rather than one you standardise on. That is not a criticism of the engine — it is a description of the information available on 10 October 2026.

Deciding

Take Qwen-Audio-3.0-TTS-Plus when cost and streaming infrastructure drive the decision. Sub-$20 per million characters, a cancellable WebSocket with 48kHz Opus output, an instruction parameter and emotion tags, and a sixth-place provider-board rank make it the most economical well-scored voice in this comparison. Take the 3.1 tier instead if you want the model that actually came second on the controlled board, and verify the voice list before assuming the swap is free.

Test HeyGen Voice when clone fidelity is the product. One speaker, many lines, a voice that must be *that* voice every time — that is the axis the controlled board measures and the axis on which it leads the entire field. Budget for a measurement period, because the credits model means you will learn your true cost by running your own text rather than by reading a rate card.

And ignore the comparison entirely if the workload is Chinese-language and real-time. Neither Elo number was measured on your voices, in your language, at your latency budget.

Keeping both reachable

This is one of the pairings where the routing layer earns its place rather than being a bolt-on. A single API key covering both vendors — provider list price passed through with no markup, so Alibaba's published $19.30 is what you pay and a Qwen price move is live the same day — removes the need to pick a winner before you have evidence. Automatic failover covers the obvious risk in a voice pipeline, where a stalled stream is audible to the listener, and the routing DSL lets you send bulk narration to the cheap engine and clone-critical lines to the one that scores 79 points higher, without two client libraries and two billing relationships. OrcaRouter's value here is not that it hosts a voice model; it is that it makes the cost of finding out which one you want roughly zero.

The open questions

Three things would settle this. A voice rate on HeyGen's own pricing page would turn the $30.00 the arena derives into a number you can audit. A HeyGen entry on the Provider Voices board would tell us whether the cloned-voice lead survives native voices — the test Qwen-Audio-3.0-TTS-Plus passes at sixth of ninety-six. And a controlled-voice result for Qwen-Audio-3.1-TTS-Plus against HeyGen Voice after more votes would tell us whether 16 Elo is a stable ordering. Until then: Qwen is the priced, streamable, well-ranked option; HeyGen Voice is the higher-scoring, unpriceable one; and the tier you benchmark matters more than the headline you read.

What this article says