How speech models fail where it matters the most and what to do about it
Together AI
Speech recognition systems achieve an average 39% transcription error rate on street names from diverse speakers, with an 18% accuracy gap between non-English and English primary speakers. The researchers reduced these errors by up to 60% using cross-lingual style transfer on fewer than 1,000 synthetic training samples. These improvements address a critical gap where street name errors in navigation and emergency dispatch systems cause significant delays and economic losses, particularly affecting non-English speakers.
Why it matters
State-of-the-art speech models like Whisper and Deepgram score near-human on benchmarks — then fail 39% of the time on street names. New research from Together AI exposes the gap and a fix.