Cartesia releases Sonic-3.6 text-to-speech model, achieving top rankings on Artificial Analysis benchmarks
Model release Provisional 95% confidence first seen
Cartesia released Sonic-3.6, a streaming text-to-speech model built on state space models that achieved first-place rankings on Artificial Analysis speech leaderboards with sub-90ms latency to first audio. The model supports 44 languages and is available as a hosted API starting at $5 per month, priced at $49 per million characters, undercutting competitors like ElevenLabs.