Improved Gemini audio models for powerful voice experiences
Google DeepMind
Google released an updated Gemini 2.5 Flash Native Audio model that improves voice agents' ability to handle complex workflows, follow instructions, and maintain natural conversations. The model achieved a 71.5% score on ComplexFuncBench Audio for multi-step function calling and reached a 90% adherence rate to developer instructions, up from 84%. The update enables live voice agents across Google products and introduces live speech-to-speech translation in Google Translate, supporting over 70 languages while preserving speaker characteristics like intonation and pacing.