TLDRocket
Sign in

Real-time speech-to-speech translation

Google Research

Google developed an end-to-end speech-to-speech translation model that translates spoken audio directly into target language audio while preserving the original speaker's voice. The system achieves a 2-second latency compared to 4-5 seconds in prior cascaded approaches, and currently supports five Latin-based language pairs including English-Spanish, English-German, and English-French. The technology has been deployed in Google Meet on servers and as an on-device feature on Pixel 10 devices to enable real-time cross-language communication.

Why it matters

Algorithms & Theory

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.