mmBERT Released, Supporting 1,800 Languages
·2025.09.09 09:00
Key point
mmBERT, a next-generation multilingual encoder based on the ModernBERT architecture and supporting over 1,800 languages, has been released.
Details
mmBERT is a state-of-the-art multilingual encoder model trained on over 3T tokens across more than 1,800 languages. Built on the ModernBERT architecture, it boasts very fast inference speed and is the first multilingual model to surpass the performance of the existing XLM-R.
Key Features and Innovations:
- Progressive Language Inclusion: New languages are added while progressively flattening the language distribution across training stages. This maximizes training efficiency for low-resource languages while maintaining overall data quality.
- 3-Stage Training Process:
- Pre-training (2.3T tokens): Covers 60 languages, with a 30% mask ratio.
- Mid-training (600B tokens): Expanded to 110 languages, with context extended to 8192 tokens, at a 15% mask ratio.
- Decay phase (100B tokens): Includes all 1,833 languages, finishing training with a 5% mask ratio.
- Inverse Mask Ratio Schedule: As training progresses, the mask ratio is progressively lowered from 30% $\rightarrow$ 15% $\rightarrow$ 5%, driving efficient training.
- Architecture and Tokenizer: Inherits the structure of ModernBERT, but uses the Gemma 2 tokenizer to enhance multilingual capability.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.