AivexaNewsSearch
AI news for builders and product teamsChecked every hour

The Open ASR Leaderboard Adds Its First Global South Language

Collected Oct 1, 2026

Voice Arena and Hugging Face have partnered to add two evaluation sets to the Open ASR Leaderboard: Monsoon en-IN and Monsoon hi-IN. According to the announcement, Hindi, spoken by more than half a billion people, is the first Indic language on a multilingual tab that previously covered only European languages. The four splits are speaker-disjoint and comprise 4,888 speakers, with 12 speaker attributes recorded for each. Each set is released as a public split, available for self-scoring, and a private split withheld to limit benchmark-specific optimisation.

The sets ship 18 columns per segment, of which 12 are metadata. Data comes from unscripted dual-channel spontaneous conversations, with clips segmented from a single channel so each carries one speaker. Contributors were recruited through the Voice Arena community and recorded two-person conversations over a peer-to-peer interface on their own handsets. The Indian English public set draws on 428 native districts across 30 states and union territories; the Hindi sets span 202 and 295 districts. Recordings come from 315 to 582 distinct device models, with no single model exceeding 2.1% of segments in any subset.

Because Hindi has far more orthographic variation than English and no fixed mapping collapses it, the Hindi sets ship a lattice: for each span of the transcript, a list of spellings accepted as correct. For Hindi, the leaderboard reports the Orthographically-Informed Word Error Rate (OIWER), introduced by AI4Bharat, instead of WER. The implementation, voi-oiwer, is open sourced.

The announcement states that eight models on the leaderboard land between 4.81 and 4.99 WER on the public Indian English split. Grouping speakers by region, it says, openai/whisper-large-v3-turbo varies by 0.46 points across zones, while mistralai/Voxtral-Mini-3B-2507 varies by 1.68. Indian-English joins the main leaderboard as Voice Arena Monsoon in the default column set, and the public and private Hindi sets appear in the Multilingual tab. These four sets are part of Monsoon, Voice Arena's broader dataset initiative for the Global South.

Read at Hugging Face Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt