TLDRocket
Sign in

Google launched Gemini 3.8 Flash TTS and Flash-Lite TTS

Google ● Covered by 9 sources

Google launched two new Gemini text-to-speech models. They’re built for custom voices, dubbing, and long-form audio.

Based on reporting by Google — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Google is adding two new text-to-speech models to Gemini: 3.8 Flash TTS and 3.8 Flash-Lite TTS. The pitch is simple enough. Voice generation is moving from preset options to something closer to a creative tool, with more control over how speech sounds and behaves.

The heavier model, Gemini 3.8 Flash TTS, is aimed at character work and detailed direction. It can create new voices from scratch using prompts, and it lets users steer delivery line by line with notes about pacing, dialect shifts, and even backchanneling. Google says it can handle more than 100 languages and dialects, and it offers 2,000-plus production-ready voices.

Flash-Lite is the scale play. Google positions it for high-volume dubbing, audio creation, and voice agents, with fine control over tone and pacing. Both models are available now in Google AI Studio and the Gemini API for developers, while enterprise access is coming later through Gemini Enterprise. Gemini Notebook gets support too, and Flash-Lite is also headed to Google Vids for everyone.

Google is also leaning hard on safeguards, which makes sense given the voice-cloning angle. Replication requires verbal consent from the voice owner, and every generated clip gets a SynthID watermark. The company says the new models are already showing up well in outside evaluations too, including Hume AI’s Voice Design Benchmark and Voice Arena.

The rollout comes with some obvious partners: Agora, LiveKit, Pipecat, Vercel, Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang. That tells you what Google wants this for. Not toy demos, but production dubbing, localized media, and voice agents that have to sound less like a robot from 2019.

My take — AI-written commentary, not fact-checked reporting

Google is doing the sensible thing here: selling voice as infrastructure, not magic. The consent checks and watermarking are the least interesting part to marketers, which is exactly why they matter. Everyone wants expressive AI audio; nobody wants another sloppy remix of someone else’s voice floating around the internet.

Read more about this at: Google

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.