Kandinsky 6.0 Video gives you a 29B-parameter checkpoint you can download and a five-second clip in return. Wan 2.7 gives you four task-specific variants, clips up to fifteen seconds, and a per-second invoice. Those two sentences are the comparison, and the reason it is worth writing down is that most head-to-head pages compare them on visual quality, where neither has an independent score that would settle it — and skip the axis where they genuinely diverge, which is what each one expects you to bring.
Sber's Kandinsky Lab open-sourced Kandinsky 6.0 Video in the first week of October 2026 under MIT, in two sizes — Pro at 29B and Lite at 3B — with code, checkpoints and diffusers integration. Alibaba's Wan 2.7 shipped on 3 April 2026 as a suite: text-to-video, image-to-video, reference-to-video and video-edit, billed per second of generated video. One of those is a model you run; the other is a service you buy. The interesting question is not which is better but which one is the right shape for the work.
The one axis where they are directly comparable
Both models generate video with sound, not video that you then score. That is rarer than it sounds and it is why these two get mentioned together at all — most of the field still produces a silent MP4 and leaves the audio to a separate pass.
Kandinsky 6.0 Video does it with a dual-stream CrossDiT: a video stream initialised from the Kandinsky 5.0 Video checkpoint, an audio stream trained from scratch on large audio corpora, and bidirectional cross-attention joining them, trained jointly after a continuous pretraining stage. Output is 44 kHz audio with lip-sync, from a text prompt (T2AV) or a reference image (I2AV).
Wan 2.7 produces native audio too, without publishing the mechanism in the same detail. Its own documentation names audio quality and on-screen text accuracy as the two areas still needing work — a fidelity caveat worth holding onto, because it is the vendor saying it, not a reviewer guessing.
That is where the likeness ends. Everything below this line is a divergence, not a ranking.
Five seconds against fifteen, and what that does to a shot list
Kandinsky 6.0 Video is trained at 121 frames, 24 fps — five seconds — and the release contains no extension mode. The paper, the serving recipe merged into vLLM-Omni, and the repository's performance table are all built around that single clip length. A thirty-second piece is six generations and five joins, and the joins are yours to hide.
Wan 2.7 covers two to fifteen seconds in one generation, decided per request. For a talking-head clip, a product insert or a single camera move, that difference is not a convenience — it is the difference between one render and an assembly step. For anything built shot by shot, five seconds is a normal unit of work and the gap mostly disappears.
• Max clip — Kandinsky 6.0 Video: 5s, fixed, 121 frames at 24 fps. Wan 2.7: 2–15s, set per request.
• Resolution — Kandinsky 6.0 Video: SD, HD and Full HD (1920×1080) through a separate super-resolution model. Wan 2.7: up to 1080p.
• Reference conditioning — Kandinsky 6.0 Video: one image, applied as a masked tail frame. Wan 2.7: up to five subject images and five reference clips, ten assets per request.
• Tasks — Kandinsky 6.0 Video: T2AV and I2AV. Wan 2.7: text-to-video, image-to-video, reference-to-video and video-edit as separate variants with separate billing.
• Weights — Kandinsky 6.0 Video: MIT, downloadable, Pro and Lite. Wan 2.7: no.
Four tasks against two modes
The task count is the part of Wan 2.7 that gets undersold in comparison tables, because it is not a quality claim — it is a scope claim. Instruction-based editing of existing footage, where you describe a change and get the same shot back with that change applied, is a workload Kandinsky 6.0 Video does not attempt at all. Reference-to-video, where a supplied performance drives a generated subject, is likewise outside its two modes.
Its billing reflects the scope. Wan 2.7's text-to-video and image-to-video variants bill on output duration only — $0.10/s at 720p and $0.15/s at 1080p on the international Singapore endpoints, with Beijing and Tokyo listing $0.086/s and $0.143/s for the same work. The reference-to-video variant bills the input video as well, capped at five seconds, so a five-second reference driving a fifteen-second generation is billed for twenty seconds of work. The video-edit variant bills output only and treats the input footage as free — which is the single most favourable rate in the family, and the reason editing is where 2.7 is genuinely cheap.
Two lifecycle facts belong beside those numbers. Alibaba attached a limited-time 30% launch discount to its own Bailian and Qwen Cloud surfaces, and it ran out on 23 September 2026. Separately, Alibaba Cloud's Model Studio retirement notice (ID 118434) takes several older models offline on 10 October 2026, and wan2.7-r2v and wan2.7-image are on that list. That is variant-level rather than a shutdown of the generation — wan2.7-t2v and wan2.7-i2v remain listed with live rates — but if reference-to-video was the reason you were looking at 2.7, that route has four days left on it.
What the independent boards do and do not show
This is the section where most comparisons of these two models quietly cheat, so it is worth being precise about what exists.
Wan 2.7 has independent scores on three separate Artificial Analysis video boards, and they do not agree with each other because they measure different things:
• Text-to-video — Wan2.7-260612 (the June build) sits 13th–14th at Elo 1,030 (±9) over 5,571 votes, at $9.00/min.
• Image-to-video — Wan 2.7 (the April build) sits 12th at Elo 1,077 (±8) over 4,571 votes, at $9.00/min.
• Video editing — Wan 2.7 sits 5th at Elo 1,077 (±5) over 17,988 votes, at $16.90/min.
Read those three together and the useful conclusion is not "Wan 2.7 is twelfth at video". It is that the model is markedly stronger at editing than at generation, on boards built for each, and that quoting a single number for it is a mistake in either direction.
Kandinsky 6.0 Video has no score on any of them. Its published evidence is vendor-side and it is worth stating exactly what it is: a VABench run in which the laboratory reports Pro leading on speech quality, audio aesthetics, text-video and audio-video alignment, lip-sync accuracy, desynchronization and visual realism among the three systems evaluated — the other two being its own Lite model and LTX 2.5 — and a lab-run side-by-side human evaluation with roughly 200 or more pairwise comparisons per criterion, in which Pro is competitive with Kling 2.6 and Veo 3.1 Fast, trails MiniMax H3 and Dreamina Seedance 2.0 on visual criteria, and stays competitive on speech quality and audio-video synchronization. None of it has been reproduced outside Sber, and none of it is an Elo on a blind public board. The right reading is "promising and unaudited", not "matches Veo".
What each one costs, stated in the terms each side uses
The two pricing models cannot be reduced to one number, and pretending otherwise is how teams make a wrong call.
Wan 2.7 is a per-second rate. A fifteen-second 1080p clip is about $2.25. A thirty-second piece is two generations plus an edit — $4.50 of raw generation plus your assembly time, or less if the editing variant is doing the work, because it does not bill the input.
Kandinsky 6.0 Video has no rate at all, because there is no service to buy. What it has is a GPU-hour cost you supply. The repository's own numbers set the floor: a Pro Full HD five-second clip is about 402 seconds of working time on an H100 after warmup, 854 on an RTX 5090, and 1,247 on an RTX 4090. The distilled checkpoint — which the paper measured at parity with the full model, mean preference 51% to 49% with no criterion reaching significance — is the one you would actually serve. Anything under 72.8 GiB of peak allocation needs block offloading, and the 16 GB preset quantizes the text encoder to NF4, which the paper confirms is the only preset change that alters the clip itself.
Put plainly: Wan 2.7's cost is a line on an invoice you can forecast. Kandinsky 6.0 Video's cost is a machine you own, a render that takes minutes per five seconds, and the option — not the obligation — to fine-tune.
Where a router sits in this
OrcaRouter does not serve either Kandinsky 6.0 Video or Alibaba Wan 2.7, and neither claim is being made here. Kandinsky is yours to download and run; Wan 2.7 is reached through Alibaba's own platforms. What a pipeline built on either model needs from a router is the text layer that surrounds the render: the prompt expansion that turns a shot description into the long or short caption these models were trained on, the caption and dialogue work that feeds an image-to-video pass, and the review summaries that decide which of several generations is the keeper. Those are ordinary text calls, and they run on one OpenAI-compatible endpoint across 200-plus models with provider list prices passed through at 0% markup, so a vendor cut is live the same day rather than after a contract cycle. Automatic failover is worth more here than in a chat workload, because a caption call that fails at minute four of a Q&A session should reroute rather than lose the render.
The decision, in one line each
If the deliverable is a finished clip with sound and you would rather not own a GPU, Wan 2.7 is the only one of the two that sells you that, and its editing variant is the strongest thing in the family on the evidence available. Check the 10 October retirement before committing to reference-to-video.
If the deliverable is a pipeline you control, if five-second units suit your shot list, and if you want to fine-tune or run offline, Kandinsky 6.0 Video is the only one of the two that permits it — at the cost of a render that is slow on consumer hardware and a quality claim that is still entirely the lab's own. The event to watch on that side is simple: an Elo.
One asymmetry is worth naming before closing, because it will outlast both models. Wan 2.7's ceiling is contractually and technically Alibaba's to set — when the suite's variants are retired or repriced, your pipeline inherits the decision. Kandinsky 6.0 Video's ceiling is your own hardware, and the MIT licence means a fork can outlive the lab's roadmap. Teams that have been burned by a model disappearing from a catalogue tend to weight that difference more heavily a year later than they did on the day they chose.












