Google’s new speech model Gemini 3.8 Live supports real-time reasoning
SiliconANGLE Mike Wheatley ● Covered by 6 sources
Google launched Gemini 3.8 Live to cut lag in voice AI. It can think and talk at once, and even hide tool calls in the background.
Based on reporting by SiliconANGLE, Mike Wheatley — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google is pushing voice AI toward something closer to an actual conversation. With Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, the company says it is tackling the lag that has long made spoken assistants feel awkward and mechanical. The pitch is simple: talk, think, and work at the same time.
The two models are described as Google’s most advanced voice systems yet. They support near-real-time reasoning and simultaneous speech-and-thought processing, while also running third-party tool and API calls in the background. In practice, that means the assistant can keep talking naturally instead of freezing up while it goes off to fetch information or complete a task.
Google is backing that claim with benchmarks. Gemini 3.8 Live Extended Thinking reached 82.6 on the Artificial Analysis Speech to Speech Quality Index, which Google says is a new high. The standard Gemini 3.8 Live placed second on Speech Agent Arena and first on ServiceNow’s EVA-Bench. Extended Thinking also posted scores of 68.6% on T-Voice, 35.1% on T-Voice-banking and 97.7% on Big Bench Audio.
The models also come with language detection, mid-conversation language switching, support for speech in 97 languages and near real-time visual grounding. Google says they can even use little verbal cues like “let me check that” to sound more natural while they work. That’s the product idea here: less robotic waiting, more of the messy overlap that happens in human speech.
Gemini 3.8 Live is available now through the Gemini API and Google AI Studio, plus private preview access in Gemini Enterprise and Search Live. Extended Thinking is available through those same channels, and also in Google Workspace apps including Docs, Gmail and Keep for subscribers, as well as in the Gemini Live apps. Google is also opening the models to developers through partners such as Vercel, Agora, LiveKit, Pipecat, Fishjam and Vision Agents. Audio generated by the models will carry an invisible SynthID watermark, and the standard version is priced at $0.005 per minute for audio inputs and $0.018 per minute for outputs. Extended Thinking also adds charges for reasoning tokens and extra inputs like video and documents.
My take — AI-written commentary, not fact-checked reporting
Voice AI has spent years sounding like it was always one tab away from a loading spinner, so this is the right problem to attack. The interesting part isn’t the benchmark bragging; it’s Google trying to make agents feel less like software and more like someone who can keep talking while looking things up. That’s a much bigger deal than another shiny demo with perfect pronunciation.
Read more about this at: SiliconANGLE