AI Pulse by Inblix

Hugging Face and IISc aim to fix AI's India problem with 150K hours of speech

Hugging Face Blog · Feb 27, 2025 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Hugging Face and IISc aim to fix AI's India problem with 150K hours of speech

The uncomfortable truth about most AI voice assistants is that they fall apart the moment you step outside a handful of dominant languages. A new partnership between Hugging Face and the Indian Institute of Science (IISc) and its ARTPARK initiative takes direct aim at that failure by supercharging Project Vaani, an effort to build what may become the most linguistically diverse open-source speech dataset on the planet.

Vaani isn’t chasing the same ten languages everyone else is. Launched with Google in 2022, the project has already open-sourced data from 80 districts, and Phase 2 is now underway across 100 more. The geo-centric collection method is what makes it genuinely different — researchers aren’t just parking themselves in Mumbai and Delhi. They’re capturing dialects from remote regions that commercial models have ignored entirely. The numbers are staggering: targets of 150,000 hours of speech and 15,000 hours of transcribed text from a million people across all 773 districts.

As of February 2025, the open-sourced portion already covers 54 languages with transcribed audio from roughly 700,000 speakers. That’s not just a big number on a slide deck. For engineers, it means you can actually train speech-to-text models that handle code-switching between Indic languages and English — the way real people actually talk. The dataset includes smaller, segmented audio units matched with precise transcriptions, which opens the door to speaker identification models, language identification systems, and even speech enhancement tools that understand the acoustic environments of Indian villages, not just soundproofed recording studios.

What makes this partnership timely is the current LLM moment. Everyone is racing to build multimodal models, but those models are starved for training data that reflects how most of the world speaks. Vaani’s spontaneous speech — collected in real-life settings, from people with wildly different educational and socioeconomic backgrounds — is the kind of resource that could shift foundational speech models from being excellent at English and French to being genuinely capable in Kannada, Bhojpuri, or Khasi. The question now is whether developers will actually use it. The dataset is there. The pipeline for real-world impact — telehealth, voter helplines, local media — is obvious to anyone paying attention. The gap between open data and shipped products, however, remains the hard part.

💡 Key Takeaways

  1. Project Vaani targets 150,000 hours of speech from 1 million people across all 773 Indian districts, dwarfing most existing multilingual datasets.
  2. The dataset prioritizes dialects from remote regions using a geo-centric approach, not just the dominant languages that commercial models already support.
  3. Vaani's transcribed subset of 790 hours from roughly 700,000 speakers can train models for code-switching, a feature most Indic speech systems currently lack.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles