Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas
MarkTechPost 1 week ago 51 ● 2 sources
Cartesia released Sonic-3.6, a real-time text-to-speech model built on state space models rather than transformers. The model ranks first on both Artificial Analysis speech leaderboards with 1,283 Elo on Provider Voice and 1,123 on Controlled Voice, and claims sub-90ms latency to first audio. It is available as a hosted API starting at $5 per month, positioning it at $49 per million characters, half the price of ElevenLabs Eleven v3.