TLDRocket
Sign in

HuggingFace, IISc partner to supercharge model building on India's diverse languages

Hugging Face Blog

The Indian Institute of Science and Hugging Face partnered to make Vaani, an open-source dataset covering India's languages and dialects, more accessible to global AI developers. Vaani targets collection of over 150,000 hours of speech data from 1 million people across all 773 districts, with Phase 1 covering 80 districts already open-sourced and Phase 2 expanding to 100 additional districts. The partnership aims to enable development of AI systems for speech recognition, language identification, and multilingual applications tailored to India's linguistic diversity.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.