HuggingFace, IISc partner to supercharge model building on India's diverse languages
Hugging Face
IISc and ARTPARK are teaming up with Hugging Face to open up Vaani, a massive Indian-language speech dataset. It could make AI actually work for the hundreds of languages big tech usually ignores.
Based on reporting by Hugging Face — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
India has more than 700 districts, dozens of major languages, and thousands of dialects that rarely show up in any AI training set. Project Vaani, started back in 2022 by IISc, ARTPARK, and Google, has been quietly trying to fix that by recording ordinary people talking in their own tongues, in their own towns, rather than harvesting whatever polished audio happens to exist online. Now Hugging Face is stepping in to host and distribute the dataset, which should make it far easier for developers anywhere to actually use it.
The numbers give a sense of the ambition here. The project is aiming for 150,000 hours of speech and 15,000 hours of transcribed text from a million speakers across all 773 districts in the country. Phase 1, covering 80 districts, is already open-sourced. Phase 2, now underway, adds another 100 districts, meaning Vaani will soon have a footprint in every single Indian state. As of mid-February, there's also a smaller transcribed subset available: 790 hours of audio from roughly 700,000 speakers, paired with 70,000 images, small enough to work with directly for training speech recognition or language models without wading through untranscribed files.
What makes Vaani genuinely different from the usual speech corpora is its geographic method. Instead of chasing the biggest languages first, researchers went district by district, capturing spontaneous, real-world speech rather than scripted studio recordings. That's the kind of messy, accented, code-switched audio that actually reflects how 1.4 billion people talk, and it's exactly what's missing from most existing datasets covering Hindi, Tamil, or Bengali, let alone the smaller regional languages.
The practical uses read like a checklist for anyone trying to build AI for the next billion users: speech-to-text and text-to-speech systems, code-switching ASR that handles Indic-English mixing, speaker verification models trained on over 80,000 voices, language identification tools, and speech enhancement systems. IISc and ARTPARK are positioning this as groundwork for things like telemedicine bots, voter helplines, and multilingual smart devices, applications that only work if the underlying model actually understands the person talking to it.
Hugging Face's role is mostly plumbing, but plumbing matters. A dataset locked in an academic repository with clunky access rarely gets used outside a small circle of researchers. Putting Vaani on Hugging Face means any developer, in India or anywhere else, can pull it into a training pipeline without jumping through institutional hoops, which is often the actual bottleneck for progress on underrepresented languages, not a shortage of research interest.
My take — AI-written commentary, not fact-checked reporting
This is the sort of unglamorous infrastructure work that actually moves the needle on AI equity, way more than another benchmark paper about multilingual capability gaps. Western labs keep training on the same handful of languages and calling it global coverage, so a geo-first dataset built by people who actually live in the districts they're recording is worth ten papers about theoretical inclusivity. My only worry is scale versus stamina — India has done ambitious open-data projects before that stalled at Phase 2, so I'll believe the full 773-district coverage when I see it on the map, not just in a roadmap slide.
Read more about this at: Hugging Face