SyncAI.news, a Varaisys broadcasting
HuggingFace, IISc partner to supercharge model building on India's diverse languages
HF

Hugging Face Blog

· 1 min read

AI LabsHugging Face Blog

HuggingFace, IISc partner to supercharge model building on India's diverse languages

The Indian Institute of Science IISc and ARTPARK partner with Hugging Face to enable developers across the globe to access Vaani, India's most diverse open-source, multi-modal, multi-lingual dataset. Both organisations share a commitment to building inclusive, accessible, and state-of-the-art AI technologies that honor linguistic and cultural diversity.

Partnership

The partnership between Hugging Face and IISc/ARTPARK aims to increase the accessibility and improve usability of the Vaani dataset, encouraging the development of AI systems that better understand India's diverse languages and cater to the digital needs of its people.

About Vaani Dataset

Launched in 2022 by IISc/ARTPARK and Google, Project Vaani is a pioneering initiative aimed at creating an open-source multi-modal dataset that truly represents India's linguistic diversity. This dataset is unique in its geo-centric approach, allowing for the collection of dialects and languages spoken in remote regions rather than focusing solely on mainstream languages.

Vaani targets the collection of over 150,000 hours of speech and 15,000 hours of transcribed text data from 1 million people across all 773 districts, ensuring diversity in language, dialects, and demographics.

The dataset is being built in phases, with Phase 1 covering 80 districts, which has already been open-sourced. Phase 2 is currently underway, expanding the dataset to 100 more districts, further strengthening Vaani's reach and impact across India's diverse linguistic landscape.

Key Highlights of the Vaani data set, open sourced so far: (as of 15-02-2025)

District wise language distribution

Transcribed subset

  • Speech Recognition: Training models to accurately transcribe spoken language.
  • Language Modeling: Building more refined language models.
  • Segmentation Tasks: Identifying distinct speech units for improved transcription accuracy.

Utility of Vaani in the Age of LLMs

What's next

How You Can Contribute

Made with ❤️ for India's linguistic diversity

Original source

This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on huggingface.co

Similar News