HeyGen Voice vs ElevenLabs: A New #1 on Cloned Voices, and the Vendor That Still Owns the Board

Source: OrcaRouter•

HeyGen Voice vs ElevenLabs: A New #1 on Cloned Voices, and the Vendor That Still Owns the Board

HeyGen Voice took the cloned-voice arena at 1,202 Elo; Eleven v4 Turbo is third at 1,167 yet leads the 96-model provider board at 1,328.

HeyGen Voice took the number-one position on Artificial Analysis's Controlled Voice arena on 9 October 2026 with 1,202 Elo, and the model it displaced from that spot was an ElevenLabs tier: Eleven v4 Turbo, now third at 1,167. That is the headline, and it is accurate. It is also a much smaller statement than it sounds, because the same vendor's models sit second, third, fourth, eleventh and fifteenth on that same controlled board, and Eleven v4 Turbo remains first on the far larger Provider Voices arena at 1,328 Elo — 126 points clear of HeyGen Voice, which is not ranked there at all. One engine won one board. The vendor still owns the category.

The same vendor is now five of the top fifteen

The controlled board holds the voice constant — eight cloned reference voices, four US and four UK — so every entry is judged on how the engine handles a voice it did not get to choose. That board now reads: HeyGen Voice 1,202, Qwen-Audio-3.1-TTS-Plus 1,186, Eleven v4 Turbo 1,167, Eleven v4 1,160, Sonic 3.6 1,143, Realtime TTS-2 1,143. Further down, Eleven v3 sits at 1,072 and v3 Conversational at 1,056.

Five of the top fifteen entries in that field carry the same vendor name. The rival with the single best model on the board has five good ones, across two generations, at two price points, plus a conversational variant that is a different product entirely. That is what an incumbent looks like from the inside of a leaderboard lead, and it is the reason a #1 debut is worth a headline rather than a migration plan.

What the ElevenLabs models actually are

Eleven v4 Turbo is described in the vendor's own documentation as the state-of-the-art model for real-time speech synthesis, recommended for support agents, AI assistants and interactive characters, reached through the Text to Dialogue websocket under the model identifier eleven_v4_turbo. ElevenLabs states a median inference latency of roughly 100ms — vendor-reported, and the same page's footnote excludes application and network latency from that figure, so it is a model-engine number, not a round-trip promise.

Language coverage for the v4 family is the widest in this comparison by a distance: 90+ languages, from Afrikaans and Arabic through Cantonese, Hindi, Japanese, Mandarin and Zulu. HeyGen's own documentation lists 300+ pre-built voices across "dozens of languages" without publishing a count, which is a different way of describing a catalogue and not directly comparable to a language figure.

Then there is the part that does not show up on a leaderboard at all: the voice library, the stability and similarity controls, per-request tuning, and a decade of integration surface. A model comparison is not a platform comparison, and the ElevenLabs entry competes on both.

The scoreboard

• Controlled Voice Elo — HeyGen Voice: 1,202 (#1). Eleven v4 Turbo: 1,167 (#3). Eleven v4: 1,160 (#4).

• Provider Voices Elo — Eleven v4 Turbo: 1,328 (#1 of 96). HeyGen Voice: not ranked.

• Published price per 1M characters — Eleven v4 Turbo: $40.00. Eleven v4: $80.00. HeyGen Voice: $30.00 on the arena's own conversion of HeyGen's credit pricing; HeyGen publishes no voice rate itself.

• Stated latency — Eleven v4 Turbo: ~100ms median inference, excluding application and network latency (vendor-reported). HeyGen Voice: no figure published.

• Language coverage — ElevenLabs v4 family: 90+ languages (vendor docs). HeyGen Voice: "dozens of languages", count not published.

• Models in the controlled top fifteen — ElevenLabs: five tiers. HeyGen: one.

The gap between the two boards is the interesting number

Eleven v4 Turbo is 161 points ahead of HeyGen Voice on the Provider Voices board and 35 points behind it on the controlled board. Both numbers come from the same site, the same blind-listener method, and the same period. What separates them is the voice.

On the provider board, each vendor supplies its own native voices, so a company that has spent years building and curating a voice library gets to field its best. On the controlled board, that advantage disappears — the engine has to render a cloned voice it did not select. The 196-point swing between the two results is, in effect, a measurement of how much of ElevenLabs' standing comes from the voices rather than the engine. It is a large fraction. It is also not zero: 1,167 is still third of forty-two.

For a buyer, that maps onto a concrete question. If you are choosing a voice from a vendor's library, the provider board is your board, and ElevenLabs' position there is unchallenged. If you are cloning a specific speaker, the controlled board is your board, and this week it has a new name at the top.

The cost of the incumbent has not changed

Eleven v4 Turbo is $40 per million characters on the vendor's API, and the older Eleven v4 is $80 — double, for a model that scores seven Elo points lower on the controlled board. That intra-vendor spread is worth pausing on: two models from one company, 2× apart in price, three places apart in rank.

HeyGen's side of the price comparison exists, but not in a form HeyGen has published. The arena lists it at $30.00 per million characters — ten dollars under the ElevenLabs tier it displaced, and 62% of the older Eleven v4's $80 — but that figure is Artificial Analysis's conversion, not a vendor quote: HeyGen's public pricing page sells credits (600 for $29/month, 1,000 for $49, 1,500 for $149) and documents that consumption varies by model, duration and complexity without stating a voice rate. So the challenger with the better cloned-voice score is priced a quarter below the incumbent on a number the incumbent did not have to derive.

Where each side actually wins

Choose ElevenLabs when breadth is the constraint. Ninety-plus languages, five ranked tiers on the controlled board, a mature voice library, and a conversational variant priced at $50 per million characters that the controlled board ranks fifteenth — that combination is difficult to assemble elsewhere, and it is why the vendor's lead survived losing the #1 seat.

Test HeyGen Voice when the constraint is sameness. One speaker, thousands of lines, a clone that must sound like that speaker every time — that is the exact workload the controlled arena models, and HeyGen Voice is currently the best-scoring engine at it. The professional-clone path trained on twenty-plus minutes of audio is the relevant product for that job, not the instant clone from a single recording.

Do not read the 1,202 as evidence that HeyGen Voice beats ElevenLabs generally. It beats one ElevenLabs tier on one board. Three other tiers are within 42 points of it, and the leader of the other board is 126 points ahead of it.

Testing a challenger without a migration

The cost asymmetry here is the useful part: a voice swap is cheap to build and instantly audible if it goes wrong, and the incumbent at $40 per million characters is an expensive thing to leave running on traffic the challenger handles better. The way to resolve that is to run both behind one integration and let your own audio decide — OrcaRouter puts both on a single API key at provider list price with no markup, so a vendor rate change is reflected at source rather than at contract renewal, and automatic failover means a voice request that fails falls through to the other engine instead of going silent. For a first pass you can weight traffic toward whichever engine your ear prefers on a small fixture, and the routing DSL lets you hold one engine for pre-rendered narration and the other for interactive turns without a second client library.

What would change this story

HeyGen Voice entering the Provider Voices arena would tell us whether a cloned-voice lead translates into competitive native voices — the test ElevenLabs passes at 1,328. A voice rate on HeyGen's own pricing page would turn the arena's derived $30.00 into arithmetic you can audit. And another month of votes would show whether 16 Elo over Qwen-Audio-3.1-TTS-Plus is a stable ordering or a snapshot. Any of the three could move the recommendation; none of them has happened yet, and the honest summary until they do is that HeyGen Voice won the controlled board and ElevenLabs still runs the category.

What this article says