TLDRocket
Sign in

MiniMax Speech 2.6 Turbo now available natively on Together AI

Together AI

Together AI now hosts MiniMax's Speech 2.6 Turbo, a top-ranked text-to-speech model, on dedicated servers. It means natural-sounding, emotional AI voices in 40+ languages can now run with sub-250ms latency alongside your chatbot infrastructure.

Based on reporting by Together AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Voice AI has had an annoying tradeoff baked into it for years: you could get a voice that sounds genuinely human, or one that responds fast enough for a real conversation, but rarely both from the same vendor. Together AI's answer, announced today, is to put MiniMax's Speech 2.6 Turbo on dedicated GPU infrastructure that sits right next to the LLM and speech-recognition workloads customers are already running. No more stitching together three vendors just to launch a voice agent that sounds decent and doesn't lag.

The numbers here matter more than the marketing copy suggests. Speech 2.6 Turbo currently sits at the top of the Artificial Analysis Arena, a blind human-evaluation leaderboard for TTS models, and it was trained on conversational data from Talkie, MiniMax's companion-chat app that reportedly has 150 million users averaging sessions over 90 minutes. That's a very different training diet than the audiobook and podcast narration most TTS models learn from, and it shows up in how the model handles prosody and pacing during back-and-forth dialogue rather than scripted reads.

A few features stand out. The model can clone a voice from just 10 seconds of audio and then speak that voice in more than 40 languages, handling messy source recordings — background noise, accents, stumbles — without falling apart. It also picks up emotional tone automatically: feed it apologetic language from an LLM and it shifts to a softer, empathetic delivery; feed it a warning and it gets serious. No SSML tags, no prompt hacking required. And it can switch languages mid-sentence in real time, detecting the boundary and applying native pronunciation, which is the kind of thing that sounds like a demo trick until you realize how many global support teams actually need exactly that.

Latency lands under 250 milliseconds on Together's dedicated endpoints, which the company frames as the real selling point — not the voice quality alone, but voice quality without the network hop tax of routing through a separate TTS vendor. Together is positioning this alongside its existing lineup of Orpheus and Kokoro for cheap high-volume jobs and Rime Arcana v2 for deterministic enterprise pronunciation, slotting MiniMax in as the premium, most-expressive tier. It's also pitching compliance boxes — SOC 2 Type II, HIPAA readiness, zero data retention — aimed squarely at enterprises nervous about routing voice data through yet another third party.

Whether this actually collapses the multi-vendor patchwork problem depends on adoption, obviously, but the logic is sound: if your speech recognition, reasoning, and voice synthesis all run on the same rails, you stop debugging latency spikes by guessing which vendor's network hiccupped this week.

My take — AI-written commentary, not fact-checked reporting

I'll believe the

Read more about this at: Together AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.