The Open ASR Leaderboard Adds Its First Global South Language
Hugging Face
Hugging Face and Voice Arena added Hindi and Indian English to the Open ASR Leaderboard. It’s the first Global South language on a board that can now show who a model fails, not just how much.
Based on reporting by Hugging Face — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Hugging Face and Voice Arena have brought Hindi and Indian English into the Open ASR Leaderboard with two new evaluation sets: Monsoon en-IN and Monsoon hi-IN. It’s a small release in hours and a big one in reach. Hindi, spoken by more than half a billion people, is the first Indic language on a multilingual tab that had been European-language-only until now.
The point of Monsoon is not just to add another score. It is to make the score less slippery, and more honest about who it serves. Each language ships with public and private splits, speaker-disjoint, and the whole collection covers 4,888 speakers with 12 speaker attributes recorded per speaker. The Hindi sets use lattices of accepted spellings because a normaliser can’t flatten Hindi variants into one fixed form. The Indian English sets use standard string references.
The collection was built to vary along nine axes: geography, age, gender, vocabulary, devices, acoustic environments, speech type, speech rate, and multiple valid transcripts. That design shows up in the numbers. The English public split spans 428 districts across 24 states and 6 union territories, while the Hindi sets still cover 202 and 295 districts. No single device model dominates any subset, and more than half of the speakers contribute exactly one segment. These are short clips from spontaneous two-person conversations, not long monologues from a small cast.
And the metadata is doing real work here. Each clip carries fields like occupation, education, marital status, income band, handset brand, current city and years in the district. That makes region-level and device-level analysis possible on a public leaderboard set, instead of leaving those questions to closed benchmarks. The paper cites prior Indian ASR work that found district-level error rates stretching from roughly 4% to 44%; Monsoon is meant to let that kind of breakdown happen in the open.
The quality control is pretty strict. Contributors were screened for language proficiency, compensated, and gave informed consent. Recordings were checked for language, speaker gender, genuine conversation, signal quality and playback abuse before transcription. Reference transcripts were then built by native-speaking linguists in a five-level process, with a separate verifier after each correction round. No model that appears on the leaderboard helped write the references it will now be judged against.
My take — AI-written commentary, not fact-checked reporting
This is the right kind of leaderboard move: less cheerleading, more exposure. A single WER number is fine for a press release and lazy for everything else, so giving Hindi and Indian English their own messy, varied test sets is a useful shove toward reality. The industry keeps acting like coverage is the same thing as competence; Monsoon at least makes that excuse harder to sell.
Read more about this at: Hugging Face