Mercury Voice targets sub-500 ms agent replies
inceptionlabs.ai
Mercury Voice is now open to enterprise users. It’s built to answer voice agents in under 500 ms without giving up reasoning.
Based on reporting by inceptionlabs.ai — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Mercury Voice is now generally available for enterprise customers, and Inception is pitching it as the answer to a very specific voice problem: how do you keep an agent smart enough to reason, use tools, and follow long prompts without making callers sit through an awkward pause?
The company says the model, a diffusion LLM tuned for voice agents, gets its first answer token out in under 320 milliseconds at the median on real customer-service prompts. That matters because voice systems are supposed to stay under roughly 500 milliseconds after someone stops talking. Mercury Voice’s p95 is 750 milliseconds, which Inception says is still faster than the median of every model except GPT-OSS-120B running on Cerebras in low-effort mode.
The claim isn’t just speed. Inception says Mercury Voice beats models including GPT-6 Luna, Gemma 4 31B, GLM-5.3-Flash, Qwen3.5-397B, and Gemini 3.5 Flash-Lite on a composite of agentic and conversational benchmarks. The benchmark mix includes τ³-bench Telecom, τ³-bench Retail, τ³-bench Airline, IFBench, and BFCL v4. On those tests, Inception says the model is also about twice as fast as the next fastest one.
That’s the pitch: don’t choose between a model that reasons and one that feels instant. Mercury Voice offers low, medium, and high reasoning settings, a 128K-token context window, up to 50K output tokens, and pricing of $0.40 per million input tokens and $1.50 per million output tokens. At launch, those prices are cut in half. Inception says a typical voice-agent setup comes out to about $0.009 per minute of conversation, versus about $0.045 per minute for GPT-4.1.
The company is already pointing to enterprise users. Audivi AI says it uses Mercury Voice for automated drive-thru ordering. Altur says it uses it for financial-institution calls. OpenCall says it got median response latency close to 170 milliseconds on its production workload. Mercury Voice is available through the Inception API as an OpenAI-compatible endpoint, with support for LiveKit, Pipecat, Vapi, Retell, or a custom stack.
My take — AI-written commentary, not fact-checked reporting
The voice-agent market has spent way too long pretending latency is a side issue. It isn’t. If a model can sound smart and still answer before the conversation gets weird, that’s the whole game — and the rest is slide-deck perfume. The only thing more predictable than AI hype is companies rediscovering that humans hate waiting on hold.
Read more about this at: inceptionlabs.ai