Let’s be honest, there isn’t a one-size-fits-all platform for developing voice agents. And to make matters a bit more blurry, if you’re an engineer comparing Retell, Vapi, and LiveKit, each is going to easily handle the standard voice AI demo.
We believe the differences between each platform begin to show themselves after the demo: when someone goes to iterate on the agent or once the agent is deployed at production scale.
Instead of doing a giant feature comparison between each, we’ve attempted to simplify the comparison into four key questions for engineering teams:
- Will engineers keep iterating on the agent after launch, or does it need to be iterated and updated without them?
- What happens when a requirement or optimization falls outside the defaults?
- Will the agent need to work beyond phone calls, on web, mobile, video, devices or WhatsApp?
- Is cost at scale a relevant concern?
Quick comparison#
We’ve included a high-level table below to outline some of the differences between Retell, Vapi, and LiveKit, but we suggest you dive into the details to understand the nuance.
* Platform, telephony and observability fees on public non-enterprise list prices, September 2026, with your own SIP provider. Excludes model costs and carrier charges, which are roughly similar on all three. Skip to question 4 for more detailed numbers and assumptions.
Four questions to help you evaluate Retell vs. Vapi vs. LiveKit#
1. Will engineers keep iterating on it after launch, or does it need to run without them?#
Short answer
Retell is built so a non-engineer can improve the agent, which works until the change isn't a field Retell exposes. Vapi lets a developer script changes while Vapi runs the call, which means developers live inside Vapi's schema and versioning. LiveKit’s open source Agent framework puts the agent in your repo, which gives maximum flexibility for iterations and optimizations, but assumes you have engineers to keep iterating on it.
The reality is that, even for seemingly straightforward tasks, there's no such thing as the perfect AI voice agent out of the box. Any use beyond the demo will surface things the demo didn't such as:
- A caller who pauses mid-sentence and gets cut off
- A tool that takes three seconds at p95 and leaves dead air
- A voicemail that gets a full sales pitch
- A transfer to a human that drops the conversation context.
If you want to optimize your agent to achieve its goal, every agent you build with any platform will inevitably need to be tweaked in month two. The question is who makes them and what a change looks like.
2. What happens when a requirement or optimization falls outside the defaults?#
Short answer
The difference comes in how developers make a change and their limits. On Retell, you stop where the slider or field stops. On Vapi, you PATCH the assistant until the schema runs out. On LiveKit, it's a code change to the agent itself in your repo. The hosted platforms ship more out of the box, and if your workflow fits perfectly within their presets you might launch faster. However, many agents outgrow presets, and that's usually where working in code pays off. Instead of hunting for a setting that might do what you need, you write the change, test it like the rest of your code, and ship it using your git-based workflows (often with a coding agent doing much of the work).
Defaults and packaged integrations help teams get standard workflows running quickly. In some cases, though, they come at the cost of developer experience: the engineers responsible for improving the agent end up working inside a “black box.”
For basic scheduling, FAQ and lead-qualification agents, defaults are often the right trade, and both Retell and Vapi add packaged helpers like Retell's managed knowledge base and eleven first-party connectors for CRMs, calendars, helpdesks and knowledge sources or Vapi's “Campaigns” feature. The question is what happens when a requirement sits half a step outside the default workflow. In those instances, developers (and their coding agents) reap the benefits of a fully flexible open source framework like LiveKit’s.
Below is a non-comprehensive list meant to illustrate examples of what we mean by “configuration” and “defaults” versus “flexibility”.
Compare the code: five real-world scenarios#
The table above outlines some examples of defaults vs. flexibility of the platforms, but developers know push comes to shove when you need to actually implement, read, and iterate the code. To help developers understand in more detail what it’s like to work with each platform, we’ve outlined a few of the scenarios below with some example code.
Scenario 1: A caller pauses mid-sentence and the agent cuts them off
"My account number is… hang on…" and the agent is already answering a question the caller hadn't finished. Every platform's defaults are tuned for a caller who talks in one steady stream. Fixing it for the caller who doesn't is usually the first optimization a team makes, and a good test of how far each platform's knobs reach.
Retell gives you a Response Wait time slider (0 to 5.5 seconds; responsiveness, 0 to 1, in the API), interruption_sensitivity, and an enable_dynamic_responsiveness toggle that adapts to the caller's pace. Retell already waits longer when it thinks the caller hasn't finished, but you can't choose or tune that detection. Switch STT to custom mode and you can also set the provider's endpointing_ms. To wait longer only after the agent asks for a number, build a Conversation Flow and override those settings on the node that collects it.
Vapi exposes far more in startSpeakingPlan: pick a smart endpointing provider (Vapi's docs recommend LiveKit's turn detector for English) or, without one, set separate wait times for turns that end on punctuation, no punctuation, or a number, and add regex rules that override either so the assistant waits longer after it says "account number." stopSpeakingPlan tunes interruption by word count and seconds. It's a PATCH to the assistant or a dashboard edit, and for this scenario it's enough. The ceiling is that the rule has to be expressible as a regex and a number of seconds.
LiveKit treats turn-taking as a pipeline stage you configure in code: choose the detector (the turn-detector model, VAD, STT endpointing, or manual), set min_delay and max_delay, switch endpointing to dynamic so the delay tracks the caller's actual pauses, and use adaptive interruption so an "uh-huh" doesn't stop the agent. Because it's code, the rule can be anything your app knows. Here the agent that collects the number changes the timing on entry, and it takes effect on the caller's next turn:
Scenario 2: A caller reads out an order number and the transcript comes back garbled
The caller says "K7 dash 3R9Q." The transcript says "k7 3 are 9 queue." What happens next is where the platforms split.
Retell hands the transcript to the model. You can boost keywords in the STT, which helps, but the model still has to guess the order number, call your lookup tool with its guess, and ask the caller to repeat when the lookup fails. Fixing the transcript before the model sees it means taking over generation with a custom LLM server.
Vapi works the same way. Keyword boosting on Deepgram improves recognition, but the model still receives whatever the STT produced and does the guessing. Transforming the turn before the model means hosting a custom transcriber or a custom LLM server that Vapi streams through.
LiveKit edits the turn inside the agent process. You already know the caller's phone number, so you know their open orders. Fuzzy-match the garbled transcript against that list, then replace the turn with the result: "Order K7-3R9Q, shipped Tuesday, arriving Thursday." The model never guesses, never repeats a digit string, and the caller never hears "can you say that again?" No network hop, no second server.
Scenario 3: Handle whatever picks up an outbound call: a person, voicemail, a phone menu, or nothing
Leaving a message is one field on every platform. Telling a menu from a mailbox, getting through the menu, and logging why 30% of last night's dials didn't connect is where they split.
Retell runs two detectors. Voicemail can hang up or leave a message; the IVR detector's only action is hang up, and navigating the menu is a separate feature you turn on. The two are handled separately, so a call classified as IVR never triggers the voicemail action
Vapi picks a detection provider and returns one verdict, voicemail or not. However, a menu, a full mailbox and silence all read as "not voicemail," so the agent starts its pitch. Getting through a menu depends on the model deciding to call the dtmf tool.
LiveKit classifies with an STT and LLM you pick and a prompt you can edit, returns one of five verdicts (person, voicemail, IVR menu, unavailable line, or uncertain) with the transcript behind it, and in Python starts IVR navigation on its own. Each verdict is a branch in your code and ops can get five reasons a dial failed, not one.
Scenario 4: Let a human screen the transfer
All three can brief the human before connecting. The difference is what happens after the human hears the briefing.
Retell has two modes. Warm transfer speaks a one-shot whisper, then bridges; the human can't reply. Agentic warm transfer hands off to a separate screening agent you build, which talks with the human and then bridges or cancels. A cancel or timeout sends the caller back to your agent, but your agent never learns what the human said.
Vapi has eight transfer modes. Seven are one-way: two blind, four that speak a fixed message or a generated summary, and one that runs your TwiML on the specialist's leg, then bridge. For a human who can ask a question or decline, you switch to the eighth, warm-transfer-experimental, and prompt a second assistant to run the handoff; if the human declines or doesn't answer, the caller comes back to your assistant or the call ends, per your fallback plan.
LiveKit runs the handoff as one awaited task. The human lands in a private room with the transcript and three tools: connect to the caller, decline with a reason, or flag voicemail. Whatever they choose comes back to your code.
Scenario 5: Dial a list of patients tomorrow morning and report who picked up
This is the one where the hosted platforms often offer a better out-of-the-box experience.
Retell and Vapi ship it as a product. Retell's Batch Call takes a list with per-recipient variables, a start time, a call window and reserved concurrency, and shows Sent, Picked Up and Successful in the dashboard. Vapi's Campaigns take up to 10,000 contacts, a schedule window, a concurrency cap, per-contact webhooks and a pre-dial eligibility hook. Neither needs code.
As an open source framework, LiveKit Agents has no campaign product. Each dial is one dispatch against a SIP trunk you bring, and the list, the schedule, the pacing and the report are yours. It’s a real afternoon of work that the hosted platforms save you. What you get for it is no ceiling: pacing, retries and reporting are whatever you write, and concurrency is a plan tier rather than a per-line fee.
3. Will the agent need to work beyond phone calls, on web, mobile, video, devices or WhatsApp?#
Short answer
All three answer a phone. Retell adds SMS, MMS and a web chat widget as products. Vapi adds SMS through your Twilio number and a chat API. LiveKit is the only one where the model can see video (like a screen-share), a call can have more than two participants, the agent can run on hardware that isn't a phone or a browser, or answer a WhatsApp call. If the roadmap ends at phone plus text, any of the three works. If it includes web, mobile, video or devices, this question decides it.
The table is the quick version. Here’s each channel in more detail, and what it takes to support it on each platform.
Which platforms include phone numbers?#
All three, with different limits.
Retell is the only one that sells outbound-ready numbers with no telephony account. LiveKit and Vapi each include a free US number, and both are inbound-only.
Outbound on either means bringing a carrier. Vapi imports the number and handles the rest. LiveKit exposes the SIP layer: you configure trunks and dispatch rules, pick carrier and transport, and dial with CreateSIPParticipant. LiveKit is also the only one that answers WhatsApp voice calls natively, through a Cloud-only connector.
Can the agent send and receive SMS?#
All three, in different ways: Retell as a product, Vapi through Twilio, and LiveKit with a small amount of code.
Retell is the only one with SMS as a product: two-way texting on its own numbers, MMS attachments the agent can read, agent-initiated outbound, and A2P registration handled in the dashboard.
Vapi's SMS is a thin layer over a Twilio account you bring. It sets the webhook, keeps a 24-hour session per customer, and replies. It's Twilio-only, US-only, customer-initiated only, and has no MMS.
LiveKit doesn't ship SMS as a separate product, but adding it takes little code. Sending a text is a function tool that calls your carrier's API, and receiving one is a short webhook that passes the message into a text-only agent session. Because it's your code, it works with any carrier, the agent can text first or mid-call, and an MMS photo goes to the model as an image.
Can customers chat with the agent by typing on a website?#
All three, and each widget now handles voice too.
Retell's widget does text chat and browser voice calls, and Vapi's runs in either voice or text mode, one at a time. Both add a chat API. LiveKit's hosted widget is voice-first with a text pane. A text-only widget is two flags on the agent and a fork of the open-source embed starter. It’s buildable, but not prebuilt for you.
Which platforms have mobile SDKs, and are they maintained?#
All three run in a browser; only LiveKit's mobile SDKs are current.
Retell's client SDK is browser-only, though actively maintained. Vapi lists Web, iOS, Flutter, React Native and Python. The web package is current, but (as of October 2026) Flutter and React Native were last published in mid-2025, the iOS SDK has no tagged release, and the Python client hasn't shipped since 2024.
LiveKit ships ten client SDKs (web, Swift, Android, Flutter, React Native, Unity, Unity WebGL, C++, Rust, ESP32), each usually updated weekly.
Can the agent see video or images?#
Only on LiveKit.
In an STT-LLM-TTS pipeline you sample frames from the user's camera or screen into any vision-capable LLM. With Gemini Live or OpenAI Realtime, live video streams to the model directly (Python). Vapi's web SDK can start a screen share and record the call, but nothing actually reaches the model. Retell has no video plane. Its one media path is MMS during a phone call: images, audio or video the caller texts in, which the agent can describe.
Which platforms support video avatars?#
LiveKit has 16 avatar plugins.
Vapi's Tavus integration has dropped out of its documentation and survives only as credential and voice-provider types in the API. Retell has none.
Can the agent run on hardware, including a robot?#
Only LiveKit.
It runs on ESP32, embedded Linux and Nvidia Jetson, so the SDKs behind a phone agent also power a speaker, a kiosk or a wearable. For robots, LiveKit streams camera, sensor and control data in the same room as the voice agent, and LiveKit Portal syncs frames with robot state, arbitrates control between operator and policy, and bridges ROS and LeRobot.
Can more than two people be on the call?#
Only on LiveKit.
A room holds any number of humans and agents, which is how warm transfer, supervisor barge-in, and multi-agent calls work. On Retell and Vapi a call is one caller and one assistant; a second human only arrives through a transfer, and there's no documented way to add a third party and keep the agent on the line.
4. Is cost at scale a relevant concern?#
Short answer
Around 1,000 minutes a month, cost alone shouldn’t decide this; all three land in a similar range on platform costs. Above it, they split, and Retell and Vapi run 3.5–4x LiveKit's cost (platform, telephony and observability fees; excludes model costs and carrier charges, which are similar on all three). LiveKit also gives you two levers the others don't: cut the LLM bill to a fraction of a cent with fast open models like Gemma 4 through LiveKit Inference, or self-host the whole stack and pay no platform fee at all.
At the end of the day, most teams want a measurable ROI from their investment in voice. In many cases, the biggest return comes from two things: whether your agent delivers a successful experience to your end users, and whether the platform makes it easier for your developers to build, iterate, deploy, and keep optimizing it, which saves engineering time. At scale, though, cost starts to weigh more heavily, and it’s where the platforms diverge most.
Every voice agent bill has the same four parts: a platform fee for running the agent, model costs for STT, LLM and TTS (or a speech-to-speech model), telephony, and whatever you pay for logs and recordings. The model and carrier costs are roughly the same wherever you run them. The platform fee is where the vendors differ, and it's the part that scales with your minutes.
What you pay the platform#
Below are public list prices as of September 2026. Each assumes you bring your own SIP provider (Twilio, Telnyx, etc.). Totals only sum public plan fees, per-minute session fees, platform SIP fees and observability fees. Excludes model costs (STT, LLM, TTS; Retell bundles STT in its session fee), carrier charges, additional concurrency and enterprise discounts, which are similar or negotiable on all three.
Why the gap opens up#
Retell and Vapi charge a flat per-minute platform fee. It's the same rate at your first thousand minutes and your millionth, so the bill grows in a straight line with usage.
LiveKit Cloud charges a lower per-minute rate plus a plan fee, and each plan includes a block of minutes before per-minute charges start. At low volume the plan fee makes LiveKit cost about the same as the others. As usage grows, the included minutes and the lower rate pull the per-minute cost down, which is why the gap widens with every order of magnitude. Past the point where any vendor's list price applies, all three negotiate, and the more relevant LiveKit option becomes self-hosting, where the platform fee is zero.
Model costs: same providers, different billing#
Bring the same STT, LLM and TTS to each platform and the per-minute cost is similar at launch. It diverges as the agent matures, because prompts grow: more instructions, more tools, a knowledge base.
Retell prices each model per minute and scales the bill up once your prompt passes a token threshold, counting instructions, tool definitions, the transcript so far and knowledge base content. Vapi passes through the provider's per-token price. LiveKit Inference bills per token, charges repeated prompt context at the provider's cached rate, and serves fast open-weight models like Gemma 4 for a fraction of a cent per minute. Switching models is one line of code.
Concurrency#
All three include a concurrency allowance and charge for more. Retell and Vapi sell additional capacity per line per month, and with burst turned on, Retell adds a $0.10/min surcharge to calls over your limit. LiveKit's allowance scales with plan tier and can be raised on Scale, with no per-line fee.
Compliance#
All three support HIPAA workloads, but at different price points: Retell includes a BAA on pay-as-you-go, LiveKit includes HIPAA and region pinning from the Scale plan up, and Vapi sells HIPAA mode as a monthly add-on. Vapi and LiveKit offer EU regions; Retell doesn't publish one.
Self-hosting#
Only LiveKit is open source. Run it yourself and there's no platform fee and no concurrency quota from the vendor; you pay for your own infrastructure and any Cloud-only pieces you keep, like Inference, phone numbers and observability. Retell and Vapi have no self-hosted option.
Bottom line: If volume is low and will stay low, the first three questions decide it. If you're heading past 1,000 or 10,000 minutes a month, expect your prompts to grow, or need the audio path in your own infrastructure, LiveKit's pricing is built for that, and it's the only one of the three you can run without the vendor.
Putting it together: which platform fits which team?#
There's no wrong answer here, only a wrong fit. Map your team against the four questions and one platform usually falls out. For most teams, the deciding one is the second: how often you'll need to go past the defaults.
Choose Retell if the agent needs to run without engineers. It's the only one of the three where the whole agent lives in a visual editor a non-developer can own, and it ships the most as products: numbers, SMS, batch calling, a knowledge base. The ceiling is whatever the editor exposes.
Choose Vapi if a developer prefers to assemble the agent from configuration, the defaults are good enough, and the use case follows a well-worn path: simple scheduling, qualification, campaigns. Vapi keeps adding rails for those patterns. You're opting into them rather than writing and optimizing the behavior for your own use case.
Choose LiveKit if your developers would rather build the agent in code than in configuration. With LiveKit, the agent is a program in your repo, so changes go through pull requests, code review, tests and CI like the rest of your stack, and coding agents like Claude Code, Cursor and Codex can build and change it directly. That matters most once you go past the defaults (which most production agents do). It's also the stronger fit if cost at scale matters. The tradeoff is that there's no visual editor or pre-packaged outbound calling campaign product.
One more cut that's easy to miss: Retell and Vapi are built for customer-facing calls. Their feature sets (transfers, voicemail detection, campaigns, SMS follow-up, batch dialing) are primarily contact-center features, and the unit of work is focused around a phone call with one caller and one assistant. If you're building a language-learning app, a voice coaching product, a tutoring or companionship experience, an in-game character, or a voice interface for a device, the product is the conversation itself and LiveKit is the only one of the three built to support that.
Frequently asked questions#
Which platform is fastest to set up: Retell, Vapi, or LiveKit?
All three can have a phone agent answering calls within a day. Retell is fastest for a team with no developer, since the whole agent is built in a visual editor. Vapi and LiveKit both take a developer an afternoon: Vapi via a config and a webhook, LiveKit via a ~30-line agent deployed to LiveKit Cloud. For an engineering team the difference is hours, not weeks; what separates the platforms is how quickly you can change the agent after launch.
Which platform gives you telephony out of the box?
All three. LiveKit is often assumed to require a carrier, but LiveKit Cloud provisions US local and toll-free numbers directly, includes a free local number on every plan, and answers inbound calls without a Twilio account. Outbound calls, transfers and international calls run over a SIP trunk from a provider you choose, and LiveKit is the only one of the three that answers WhatsApp voice calls natively. Retell sells outbound-ready numbers with no carrier account. Vapi's free numbers are inbound-only; outbound means importing a number from Twilio, Telnyx, Vonage or a SIP trunk.
How do you deploy, manage, and host LiveKit?
Which platform has the lowest latency?
Can you use speech-to-speech models like OpenAI Realtime or Gemini Live?
Which platform should you use if voice is the product, not a support channel?
Which is the most cost-effective at scale?
Can you self-host any of these?
Can a coding agent like Codex, Claude Code, or Cursor build on these?







