Start with the arithmetic, because it is the only place these two overlap. A six-second silent 1080p clip from Hailuo 2.3, the budget video model from April 2026, lists at a handful of cents per second — roughly $0.48 for the clip. A five-second clip with synchronized 44 kHz audio and lip-sync from Kandinsky 6.0 Video, which Sber's Kandinsky Lab open-sourced under MIT in the first week of October 2026, costs nothing per clip and about seven minutes of H100 time at Full HD. That is the whole matchup in two sentences, and the decision it implies is not "which is better" — it is which meter you would rather run: a per-second invoice, or your own hardware.
Everything else follows from that split, and most comparison pages get it backwards by treating these as the same kind of product at different price points. They are not. Hailuo 2.3 is a service with a stylist's brief and no sound; Kandinsky 6.0 Video is a two-size checkpoint family — Pro at 29B and Lite at 3B — with sound as its central architectural feature.
What the cheap one is actually good at
Hailuo 2.3's reputation is earned and specific. It handles cinematic camera motion, human motion and micro-expressions well, and it is stronger than its price suggests on stylized and anime content. It produces text-to-video and image-to-video up to 1080p. MiniMax does not market it as a photorealistic model and it is not one — it is a stylist with a good eye for movement.
Its limits are equally specific, and they are the reason this comparison is short.
• No native audio. Hailuo 2.3 is silent output, and the earlier Hailuo generation that had audio does not carry forward into it. Every clip needs sound added downstream.
• Clip ceiling of six seconds at 1080p, ten seconds at 768p, with no end-frame control — so clips do not chain cleanly into a sequence.
• Not on the current independent boards. Hailuo 2.3 does not appear on Artificial Analysis's text-to-video, image-to-video or video-editing leaderboards as they read today, so there is no public Elo to compare against anything. Older references to it placing around the high twenties on a video arena do not match the boards currently published, and should be treated as a stale snapshot rather than a current position.
Pricing is quoted in two currencies by two different parties and both are worth knowing. MiniMax's own API sells prepaid point packages, where a six-second 768p standard clip burns one point and a ten-second 768p or six-second 1080p clip burns two — with the faster Hailuo-2.3-Fast variant cutting those to roughly 0.7 and 1.1 to 1.3 points. Third-party trackers list a flatter rate in the region of $0.067 to $0.082 per second, one of them at $0.0817/s and citing "up to 1080p, 10s max per clip, silent output" and about $147 a month for three hundred six-second clips. Whichever route you take, the order of magnitude is the same and it is the cheapest generation in this category.
The problem with buying into a superseded generation
There is a lifecycle fact about Hailuo 2.3 that matters more than any spec, and it is not a criticism of the model. MiniMax has moved on. Its successor, MiniMax H3 — the Hailuo 3 generation — was released on 31 July 2026 and is a different class of product: an omni-modal model reading text, image, video and audio references as one context, generating four to fifteen second clips at 768p or 2K with native stereo audio, billed per second of output.
That is not a vendor claim. It is on the independent boards, and the placement is strong:
• Text-to-video — MiniMax H3 (768p) 4th at Elo 1,137 (±9) over 8,315 votes, at $4.80/min.
• Image-to-video — MiniMax H3 2nd at Elo 1,181 (±8) over 7,284 votes, at $7.80/min.
• Video editing — MiniMax H3 2nd at Elo 1,132 (±6) over 12,043 votes, at $7.80/min.
So the honest framing of a Hailuo 2.3 purchase in October 2026 is this: you are buying the budget tier of a family whose flagship generation is top-five in both generation modes, has the audio that 2.3 lacks, and costs perhaps five to ten times more per second. That is a defensible decision for high-volume styling work — it is exactly what a budget tier is for — but it should be a deliberate one, not a default. Anyone evaluating 2.3 today should price H3 next to it first, because the gap between them is not incremental.
Five seconds against six, and why the difference is smaller than it looks
On length, Kandinsky 6.0 Video does not win. It is trained at 121 frames and 24 fps — five seconds — and the release contains no extension mode. Hailuo 2.3 gives you six seconds at 1080p. Neither is enough for a continuous scene, and both mean a thirty-second piece is five or six generations plus the joins.
The gaps that matter are elsewhere.
• Audio — Kandinsky 6.0 Video: 44 kHz with lip-sync, produced jointly with the picture, always on. Hailuo 2.3: none, by design.
• Resolution — Kandinsky 6.0 Video: SD, HD and Full HD via a dedicated super-resolution model, x2, x4 or x2.25 upscales of five-second clips. Hailuo 2.3: up to 1080p, no SR stage.
• Conditioning — Kandinsky 6.0 Video: text-to-audio-video and image-to-audio-video, one reference image applied as a masked tail frame. Hailuo 2.3: text-to-video and image-to-video, no end-frame control.
• Weights — Kandinsky 6.0 Video: MIT, Pro and Lite, with diffusers integration and — since this morning — a vLLM-Omni serving pipeline. Hailuo 2.3: not released.
• Pricing — Hailuo 2.3: per second of output, or points against a prepaid package, with a cheaper Fast variant. Kandinsky 6.0 Video: no rate at all; you supply the GPU.
What the vendor's own evidence says, and what it cannot
Kandinsky 6.0 Video's quality case is unaudited and should be described that way. Sber reports VABench results in which Pro leads on speech quality, audio aesthetics, text-video and audio-video alignment, lip-sync accuracy, desynchronization and visual realism among the three systems evaluated — its own Lite model and LTX 2.5. It also reports a lab-run side-by-side human evaluation, roughly 200 or more pairwise comparisons per criterion, in which the model is competitive with Kling 2.6 and Veo 3.1 Fast, trails MiniMax H3 and Dreamina Seedance 2.0 on visual criteria, and stays competitive on speech quality and audio-video synchronization. There is no blind public Elo, and no external reproduction.
Hailuo 2.3's evidence is the opposite shape: widely used, priced publicly, but currently absent from the independent video boards, so it cannot be placed against a current field either. Neither model has the thing a buyer would most like — and in Hailuo's case, its own successor does. That is the asymmetry worth naming: one model's family has already been scored and placed, and the other's has not been scored at all.
What is left, then, is the audio divide, and it is a genuine divide rather than a feature-list line. Hailuo 2.3's silence simplifies one thing — there is no soundtrack to preserve when you hand a clip to a mastering pass — and complicates another: every talking-head, every product spot with narration, every clip where a sound effect has to land on an action, is now two generations plus an alignment step you own. Kandinsky 6.0 Video's whole architecture is aimed at that problem: a dual-stream CrossDiT pairing a video stream with a from-scratch audio stream under bidirectional cross-attention, with a reinforcement-learning stage the paper reports as cutting generated-speech word error rate by 47% — measured with Whisper-large-v3, one of the RL reward models, and not comparable to the separate Qwen3-ASR-1.7B word-error rate the paper reports alongside VABench.
Where a router fits, and where it does not
OrcaRouter does not serve Kandinsky 6.0 Video or Hailuo 2.3, and neither claim is made here — Kandinsky is yours to download and run, Hailuo 2.3 is reached through MiniMax's own API and points system. What a pipeline built on either one needs from a router is the text around the render: the shot list, the prompt expansion that turns a brief into the long or short caption these models were trained on, the caption and dialogue work, and the review summaries that decide which generation is kept. Those are ordinary text calls, and they run on one OpenAI-compatible endpoint across 200-plus models at provider list price with 0% markup, so a vendor price change is live on the same key the day it lands, with automatic failover so a failed call mid-render reroutes instead of losing the batch. The orchestration layer around a video model is a routing problem even when the video model is not routable — which is the situation for both of these today.
Which one to pick
Pick Hailuo 2.3 when the job is volume, the aesthetic is stylized or animated, the deliverable is silent, and the per-clip meter is the constraint that matters. It is the cheapest competent generation in the category and it is priced accordingly. Price MiniMax H3 beside it before you commit, because the flagship generation of the same family is scored top-five in both modes and has the audio 2.3 never had.
Pick Kandinsky 6.0 Video when the deliverable needs sound, when five-second units suit your shot list, when image-to-video fidelity to a supplied still is the requirement, and when owning the model matters more than renting one. Go in with the costs stated plainly: no rate, a five-second ceiling, a Full HD render that runs about 402 seconds on an H100 and 854 on an RTX 5090 for the Pro model, block offloading required below 72.8 GiB of peak allocation, and a quality claim that is entirely Sber's own until somebody outside the lab scores it.
The question underneath the two price lists is the one that decides a lot of video infrastructure this quarter. Hailuo 2.3 charges you by the second and hands you a file; Kandinsky 6.0 Video charges you a GPU and hands you a licence. For a team generating a hundred clips a month, the invoice wins on every axis that a spreadsheet measures. For a team that intends to fine-tune, run offline, or still be using the same model after the next generation ships, the download wins on the one axis a spreadsheet cannot see.












