AI Briefing
KO

Xiaomi Unveils MiMo V2.5 Pro

·2026.04.28 02:57

Key point

Xiaomi has unveiled MiMo-V2.5-Pro, a 1.02T-parameter MoE model.

1 / 2

Details

Xiaomi MiMo has unveiled MiMo-V2.5-Pro. It is an MoE language model with 1.02T total parameters and 42B active parameters, supporting up to 1M tokens of context.

  • Hybrid Attention: Mixes SWA and GA at a 6:1 ratio and uses a 128 sliding window to reduce KV cache storage.
  • 3 MTP modules: Designed to boost inference speed and RL rollout efficiency.
  • Training scale: Pretrained on 27T tokens with FP8 mixed precision and a native 32k sequence length.
  • Post-training: Went through SFT, domain-specific RL, and MOPD to strengthen agentic, coding, and long-reasoning performance.

On benchmarks, the Base model recorded MMLU 89.4, GSM8K 99.6, MATH 86.2, and SWE-Bench (AgentLess) 35.7. In long-context evaluation, while the previous V2 collapsed sharply beyond 128k, V2.5 Pro is reported to have maintained BFS 0.56 / Parents 0.92 at 512k and 0.37 / 0.62 at 1M.

A deployment guide is provided based on SGLang, and the model can be downloaded from Hugging Face.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.