Language Discrimination Improves Linguistic Learning in Multilingual Speech Models
Apple Machine Learning Research
Multilingual HuBERT-style self-supervised speech models improved linguistic learning when pretraining was made to better discriminate between languages. Phone-ABX error dropped from 11.6% in the bilingual baseline to 10.4% (monolingual: 10.8%) while lexical (sWUGGY) and prosodic performance also increased. The gains were largest when language discrimination was introduced in the first training iteration, whereas later or repeated use caused more language-wise segregation.
Why it matters
Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model’s ability to discriminate languages during pretraining reduces and, on some measures, closes this multilingual gap on continuous phonetic and higher-level linguistic measures, while preserving substantial cross-language sharing. Using a controlled English/French HuBERT setting, we test two interventions which strengthen language discrimination: an auxiliary language…