Google releases Gemini 3.5 Transcribe, a new speech-to-text model for live and non-streaming transcription
Model release ● Confirmed 93% confidence first seen
Google has released Gemini 3.5 Transcribe, an AI speech-to-text model intended for real-time and application-based voice transcription. The coverage describes improvements such as reduced transcription errors and filler-word handling, with availability via Gemini API and AI Studio and rollout across Google products and voice features. Reported performance figures include lower word error rates for both streaming and non-streaming use cases, and the service is positioned as API-based rather than providing open weights.