TLDRocket
Sign in

Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas

MarkTechPost Asif Razzaq Covered by 2 sources

Cartesia's new Sonic-3.6 speech model just topped both Artificial Analysis TTS leaderboards. It beat ElevenLabs on the test that isolates the actual engine, not just the voice library.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Cartesia just put out Sonic-3.6, a new version of its real-time text-to-speech model, about three months after Sonic-3.5 shipped. The pitch this time is naturalness, and unlike most naturalness claims, this one is checkable against outside numbers. Sonic-3.6 now sits at #1 on both Artificial Analysis speech arenas, scoring 1,283 Elo on the Provider Voice board and 1,123 on the Controlled Voice board.

The second number is the one that actually tells you something. The Controlled board clones every competing model onto the same eight reference voices, which strips away any advantage from having a bigger or better voice catalog and leaves pure synthesis quality. Sonic-3.6 wins that comparison outright, with Sonic-3.5 landing second and ElevenLabs' Eleven v3 in third. So the improvement here is architectural, not cosmetic.

That architecture is unusual for the category: Cartesia builds Sonic on state space models instead of transformers, and it frames the standard speed-versus-naturalness tradeoff as something engineering choices can dissolve rather than something teams just have to accept. The practical payoff is latency — Cartesia states sub-90ms time-to-first-audio for Sonic, and 100ms transcript latency for its companion speech-to-text model, Ink-2. Both figures are vendor-stated model latency, not full round-trip numbers, so anyone building on this should benchmark their own pipeline rather than take the spec sheet at face value.

On the production side, Sonic includes features clearly built for agent transcripts rather than audiobook narration: inline tags for things like laughter, instant voice cloning from roughly 10 seconds of sample audio, custom pronunciation dictionaries with IPA overrides, and native handling of alphanumerics so order numbers and confirmation codes read correctly without preprocessing. Cartesia's demos also show Hinglish code-switching between Hindi and English, alongside more ordinary English speech with natural pauses and filler words.

Pricing tells its own story. Artificial Analysis puts Sonic-3.6 at $49 per million characters, half of ElevenLabs' Eleven v3 at $100, though still well above Speechify's Simba 3.2 at $10 for a 1,240 Elo score. Cartesia itself sells credits rather than raw characters — its Scale tier runs $299 a month for roughly 10,667 TTS minutes and 15 concurrent requests, with voice agent calls billed separately at 6 cents a minute. And it's worth flagging that this is a closed, hosted product: no open weights, no Hugging Face repo, and Sonic-3.6 is still in beta while Cartesia's own docs list 3.5 as the stable option.

My take — AI-written commentary, not fact-checked reporting

Topping a leaderboard that clones every model onto identical voices is a genuinely harder flex than topping one where a good voice library can paper over a mediocre engine, so this result deserves more attention than the usual TTS press release. That said, calling something deployable while it's in beta, priced in proprietary credits, and still trailing its own predecessor in the docs is a stretch — enterprises buying on latency claims should measure their real round trip before betting a contact center on vendor numbers. The state space model bet is the more interesting story here, and if it keeps outperforming transformer-based rivals, that's the detail worth watching, not the Elo score of the week.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.