AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Language Discrimination Improves Linguistic Learning in Multilingual Speech Models

Collected Oct 2, 2026

Apple Machine Learning Research published work showing that strengthening a multilingual speech model's ability to discriminate languages during pretraining reduces, and on some measures closes, the gap between multilingual and monolingual models. The authors are Maureen de Seyssel, Jie Chi and Zakaria Aldeneh, with Chi and Aldeneh listed as equal contributors.

The work addresses a known limitation: multilingual self-supervised speech models can share information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. Using a controlled English/French HuBERT setting, the authors tested two interventions intended to strengthen language discrimination: an auxiliary language classifier and per-language k-means targets.

Across the interventions, continuous-feature phone discrimination error (phone-ABX, lower is better) dropped from 11.6% in the bilingual baseline to 10.4%, compared with 10.8% for the monolingual model. Lexical performance (sWUGGY, higher is better) rose from 52.1% to 56.7%, against 58.5% monolingual. Prosodic performance on the ProsAudit lexical subtask increased from 68.9% to 72.9%, against 72.6% monolingual. The authors report that substantial cross-language sharing was preserved.

Across HuBERT training stages, the strongest gains on most linguistic measures occurred when language discrimination was introduced in the first iteration. Later or repeated interventions produced smaller improvements and came with increased language-wise segregation. The authors state these results support a causal role for language discrimination in reducing the additional cost of multilingual learning.

Read at Apple Machine Learning Research

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model’s ability to discriminate languages during pretraining reduces and, on some measures, closes this multilingual gap on continuous phonetic and higher-level linguistic measures, while preserving substantial cross-language sharing. Using a controlled English/French HuBERT setting, we test two