HF Introduces Hindi ASR Evaluation
Key point
Hugging Face has released Monsoon, a speech recognition evaluation dataset for Hindi and Indian English, to support Global South languages.
Details
Hugging Face and Voice Arena have collaborated to add a new evaluation dataset, Monsoon, for Hindi (hi-IN) and Indian English (en-IN) to the Open ASR Leaderboard. This marks the first Global South languages included in the leaderboard, expanding the previously Europe-centric evaluation framework.
Existing ASR leaderboards rely on a single metric (WER) and fail to capture performance disparities based on race, gender, accent, and other factors. To address these gaps, Monsoon was designed considering 9 axes: geography, age, gender, vocabulary, device, environment, speaking style, speaking rate, and multiple acceptable answers.
Dataset Composition and Features
- Scale: Includes a total of 4,888 speakers, with 12 attributes (occupation, education level, income, etc.) recorded for each speaker.
- Splits: Divided into public and private splits for each language to prevent benchmark overfitting.
- Hindi Processing: Due to the diversity of Hindi orthographic variations, reference transcripts are provided in a lattice format containing a list of acceptable spellings instead of a single ground-truth string.
- Collection Method: Based on unscripted conversational data collected using personal mobile phones in real-world environments, rather than quiet rooms.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.