TLDRocket
Sign in

Introducing Mercury 2.5

inceptionlabs.ai

Mercury 2.5 is out, with faster and smarter results than Mercury 2. It keeps low latency and low cost, which is the whole trick for voice, search, and coding.

Based on reporting by inceptionlabs.ai — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

TLDR Dev has released Mercury 2.5, calling it its most capable production model so far. The company says it is a clear jump in quality over Mercury 2, while keeping the same low-latency, low-cost serving setup that made the earlier model useful in production.

The pitch is based less on benchmark worship and more on what customers actually broke. Since Mercury 2 launched, thousands of developers have used it, dozens of enterprises have put it into production, and usage has grown by more than an order of magnitude. That feedback, along with failure cases from real deployments, was fed back into training and evaluation. Mercury 2.5 is the first model to come out of that loop.

On paper, the company is leaning hard into scale and speed. It says Mercury 2.5 is the most capable diffusion LLM on the market and, to its knowledge, the largest diffusion language model ever trained. It claims a 40% increase in intelligence over Mercury 2, 1,107 tokens per second on widely available NVIDIA GPUs, a 260K-token context window, and pricing of $0.20 per million input tokens and $0.75 per million output tokens. At launch, those prices are cut to $0.04 and $0.15.

The interesting part is where Mercury already lives. It is being used in search agents, RAG pipelines, voice agents, and coding tools, where repeated model calls make latency and cost compound fast. OpenCall says Mercury brought median response latency close to 170 milliseconds on its production workload. Augment Code says moving context compaction to Mercury cut latency by 82%, from roughly 150 seconds to 27 seconds, while cutting cost by 90%. Tool-search summaries now return in under a second.

Alongside Mercury 2.5, TLDR Dev is also previewing Mercury Voice and Mercury Router. Mercury Voice is aimed at voice agents and is described as keeping time-to-first-token under 170 milliseconds. Mercury Router is meant to direct prompts toward the best mix of open and closed models for quality, speed, and cost. The models are available through the company’s Inception API, Baseten, and OpenRouter, with enterprise options for dedicated capacity, autoscaling, compliance controls, and configurable data retention.

My take — AI-written commentary, not fact-checked reporting

This is the rare AI launch that sounds built for actual work instead of a benchmark trophy case. Low latency, predictable cost, and messy production feedback matter more than yet another miracle demo. The industry keeps rediscovering that speed is a feature, not a footnote.

Read more about this at: inceptionlabs.ai

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.