MiMo ASR Released
Key point
Xiaomi MiMo released an 8B ASR model that is robust to multiple languages, dialects, and noisy environments.
Details
The Xiaomi MiMo team released MiMo-V2.5-ASR.
This end-to-end ASR model is designed to transcribe a wide range of content, including Chinese/English, Chinese dialects, code-switching, song lyrics, noisy environments, multi-speaker conversations, and knowledge-intensive sentences.
- Training approach: large-scale mid-training, high-quality supervised fine-tuning, and a new reinforcement-learning algorithm
- Strengths: support for dialects such as Wu, Cantonese, Hokkien, and Sichuanese, plus built-in punctuation
- Evaluation: claims SOTA across public benchmarks, and highlights top-tier performance on the Open ASR Leaderboard on challenging English benchmarks like AMI
- Deployment: released together with a Hugging Face model card, a GitHub repository, and a Gradio demo
The model has 8B params, is distributed as F32 safetensors, and the demo lets users select a language tag of Chinese / English / Auto.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.