TLDRocket
Sign in

How speech models fail where it matters the most and what to do about it

Together AI

Speech recognition systems achieve an average 39% transcription error rate on street names from diverse speakers, with an 18% accuracy gap between non-English and English primary speakers. The researchers reduced these errors by up to 60% using cross-lingual style transfer on fewer than 1,000 synthetic training samples. These improvements address a critical gap where street name errors in navigation and emergency dispatch systems cause significant delays and economic losses, particularly affecting non-English speakers.

Why it matters

State-of-the-art speech models like Whisper and Deepgram score near-human on benchmarks — then fail 39% of the time on street names. New research from Together AI exposes the gap and a fix.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.