Measuring benchmark optimization in speech recognition
Hugging Face
Speech recognition model evaluation researchers showed that open benchmarks can be “optimized” by models learning benchmark-specific transcript patterns rather than transcribing the audio faithfully. In their VoxPopuli analysis, they flagged reference errors in 40% of test clips affecting about 3% of reference words, with benchmark-optimized behavior reproducing erroneous references 18–30% of the time. Held-out and newly collected audio reduces these effects because the cues tied to the original benchmark conditions are missing, so more models switch back to audio-faithful transcripts.