TLDRocket
Sign in

Announcing the fastest inference for realtime voice AI agents

Together AI

Together AI announced expanded voice infrastructure for AI agents including streaming Whisper speech-to-text with WebSocket APIs, serverless open-source text-to-speech models (Orpheus at 187ms and Kokoro at 97ms time-to-first-byte), and new transcription capabilities with Voxtral Mini and speaker diarization. The streaming Whisper transcription completes transcripts up to 35% faster than alternatives with tuned voice activity detection for natural conversation timing. These integrated services enable developers to build voice agents with lower latency, reduced operational complexity, and consistent performance at scale.

Why it matters

Together AI launches the fastest voice AI stack: streaming Whisper STT, serverless open-source TTS (Orpheus & Kokoro), and Voxtral transcription. Sub-second latency for production voice agents.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.