Improved Gemini audio models for powerful voice experiences
Google DeepMind
Google upgraded Gemini's voice AI so it holds better conversations and now translates speech live through your headphones. It's rolling into Search Live, Gemini Live, and Google Translate starting today.
Based on reporting by Google DeepMind — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google DeepMind pushed out an update to Gemini 2.5 Flash Native Audio this week, and the headline isn't just better voice quality — it's that the model finally handles the messy parts of real conversation. Function calling reliability jumped, hitting 71.5% on ComplexFuncBench Audio, a benchmark built around multi-step tasks with constraints. Instruction adherence climbed from 84% to 90%. Those numbers sound dry, but they translate into an AI agent that actually remembers what you told it three turns ago and doesn't fumble when it needs to pull live data mid-sentence.
The model is already live in Google AI Studio, Vertex AI, and is rolling into Gemini Live. For the first time, it's also landing in Search Live, which means the naturalness Google has been chasing in voice AI is now available for something as mundane as asking Search a follow-up question out loud. Enterprise customers are already leaning on it. Shopify's Sidekick bot reportedly gets thanked by users who forget it's a machine. United Wholesale Mortgage says pairing it with their Mia assistant helped generate over 14,000 loans since May. Newo.ai claims its AI receptionists can now pick out the main speaker in a noisy room and switch languages mid-call without missing a beat.
But the more interesting move is live speech translation, landing in beta in the Google Translate app today. This isn't the clunky phrase-by-phrase translation people are used to — it's streaming speech-to-speech that tries to preserve how someone actually sounds, keeping their pacing, pitch and intonation intact rather than flattening everything into a robotic narrator voice. It covers more than 70 languages and roughly 2,000 language pairs, and it can handle two people speaking different languages back and forth, auto-switching output based on who's talking.
There's a continuous-listening mode too, aimed at travelers who just want to wear earbuds and hear the world translated into their own language without touching a settings menu. Google says it auto-detects the spoken language and filters ambient noise, which matters if you've ever tried to use a translation app on a busy street and gotten nothing but garbage. The beta starts today on Android in the US, Mexico and India, with iOS and more countries
My take — AI-written commentary, not fact-checked reporting
Live speech translation that keeps someone's actual tone and pacing intact is the kind of feature that quietly changes how people travel and do business across language barriers, and it's a far bigger deal long-term than another point of improvement on a function-calling benchmark. That said, Google rolling this out region-by-region on Android first, with iOS and the API dangling somewhere in 2026, is a reminder that
Read more about this at: Google DeepMind