Apple Research: Language Discrimination Reduces Multilingual Speech Model Gap
Key point
Interventions like auxiliary classifiers reduced phone discrimination error from 11.6% to 10.4% in English/French HuBERT models.
Details
Multilingual self-supervised speech models often underperform monolingual counterparts when sharing information across languages under matched data budgets. Apple Research demonstrates that strengthening a model's ability to discriminate languages during pretraining reduces this gap on phonetic and linguistic measures while maintaining cross-language sharing.
Interventions and Results
Using a controlled English/French HuBERT setting, researchers tested two interventions: an auxiliary language classifier and per-language k-means targets. These changes yielded significant improvements across key metrics:
- Phone discrimination error (phone-ABX) decreased from 11.6% in the bilingual baseline to 10.4% (compared to 10.8% for monolingual models).
- Lexical performance (sWUGGY) increased from 52.1% to 56.7% (compared to 58.5% for monolingual models).
- Prosodic performance (ProsAudit) rose from 68.9% to 72.9% (compared to 72.6% for monolingual models).
Timing of Interventions
The study found that the strongest gains occurred when language discrimination was introduced in the first iteration of HuBERT training. Later or repeated interventions produced smaller improvements and led to increased language-wise segregation. These results support a causal role for language discrimination in reducing the cost of multilingual learning.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.