Gemini 3.1 Flash TTS: the next generation of expressive AI speech
Google DeepMind
Google introduced Gemini 3.1 Flash TTS, a text-to-speech model that adds granular audio tags allowing developers to control vocal style, pace, and delivery through natural language commands. The model achieved an Elo score of 1,211 on the Artificial Analysis TTS leaderboard and supports over 70 languages with native multi-speaker dialogue. Developers can now export precise voice parameters as API code, enabling consistent character voices across projects while all generated audio includes SynthID watermarking for AI-content detection.
Why it matters
Our newest audio model introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.