TLDRocket
Sign in

Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design

MarkTechPost Asif Razzaq ● Covered by 9 sources

Google launched two new Gemini voice models with prompt-based voice design. One’s for creative performances, the other for cheaper scale — and both are API-only.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Google has added two new text-to-speech models to its Gemini Audio family: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The company says they are its most expressive audio generation models so far, and it has split them into two jobs instead of trying to make one model do everything badly.

Flash TTS is the showier one. Google is aiming it at character voices and creative direction, with use cases like gaming, immersive audiobooks, podcasts and interactive media. It lets developers steer acting cues, pacing, dialect changes and backchanneling line by line using natural language. Flash-Lite TTS is the practical sibling, aimed at dubbing, audio content creation and expressive voice agents, with controls tuned for tone, pacing and expressive nuance.

The bigger shift is voice design itself. Previous Gemini TTS models came with 30 original voices. This release moves to a much larger system that can generate new voices from prompts describing a role, accent and other voice traits. Google says that works across more than 100 languages and dialects, and the demo examples include a Melbourne DJ, a monotone robot and a Japanese dragon. Developers also get a library of 2,000-plus production-ready voices, with regional varieties such as Mexican Spanish, Quebec French and Scots English.

Google is also pushing persistence. Custom voices can be saved and reused with minimal drift, and the models can handle long-form audio with stable quality, pacing and timbre across hours. They can stage two-speaker conversations from one script, and they accept non-verbal cues like <laughs>, <sigh> and <gasp>, plus reaction beats such as |mhm| and |yeah| for timing and comedy.

There are safety and provenance controls layered in too. Voice replication uses a 30-second audio sample, but the sample must be the user’s own voice or one they have rights to use, and it needs a verbal consent recording from the voice owner. Every clip gets a SynthID watermark, and replicated voices also carry C2PA content credentials. Google says both models are rolling out now through the Gemini API and Google AI Studio, with enterprise API access via Gemini Enterprise coming soon.

My take — AI-written commentary, not fact-checked reporting

This is Google doing the obvious but necessary thing: making voice tools more controllable, more synthetic, and less toy-like. The open-model crowd will hate the API-only bit, but the market keeps rewarding closed systems that ship the boring parts — consent, watermarking, reuse — without making teams build a compliance shrine around them. That's the real product here, not the robot dragon.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.