Over the past few months, thousands of developers have started building with Mercury across real-time voice applications, search and retrieval pipelines, and AI coding subagents. As we’ve scaled capacity, we’ve also continued improving the model itself.
Today, we’re making Mercury easier to build with.
Every new Inception API key now includes 100 million free tokens. Enough headroom to benchmark Mercury against your current stack on production workloads.
We've increased free-tier rate limits by 10x, so you can run production-like traffic without hitting a wall.
Since launch, we’ve continued improving Mercury 2 across production workloads:
- Lower time-to-first-token
Lower time-to-first-token
- Lower end-to-end latency
Lower end-to-end latency
- More reliable tool calling
More reliable tool calling
These improvements are especially noticeable for real-time voice, search, coding agents, and multi-agent workflows, where latency compounds across every model call.
Today, dozens of AI-native companies and enterprises run Mercury 2 in production.
Mercury 2 is live on Baseten today as part of the launch of Baseten for Model Labs. If your team already builds there, you can add Mercury 2 to your stack without onboarding a new provider.
For enterprise rate limits, tighter latency budgets, SLAs, or help tuning a specific workload, contact hello@inceptionlabs.ai. Response within an hour.









