TLDRocket
Sign in

​​Speech-to-Retrieval (S2R): A new approach to voice search

Google Research

Google developed Speech-to-Retrieval (S2R), a system that maps spoken queries directly to relevant documents without converting speech to text first, avoiding transcription errors that plague traditional voice search. The S2R model uses a dual-encoder architecture with an audio encoder and document encoder trained on paired audio queries and documents, and Google evaluated it on the Simple Voice Questions dataset across 17 languages and 26 locales. The S2R system significantly outperforms traditional cascade speech-recognition-based models and is now serving users in multiple languages, with Google open-sourcing the SVQ dataset to advance the research field.

Why it matters

Machine Intelligence

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.