TLDRocket
Sign in

Adding Benchmaxxer Repellant to the Open ASR Leaderboard

Hugging Face Blog

The Open ASR Leaderboard added private datasets from Appen Inc. and DataoceanAI covering scripted and conversational speech across multiple accents to prevent models from optimizing specifically for public benchmarks. The private datasets comprise 29.7 hours of audio across Australian, Canadian, Indian, American, and British English with gender-balanced speakers. The leaderboard now offers optional toggling between public and private datasets for evaluation, with the default ranking computed only on public data to preserve benchmark integrity.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.