AI Briefing
KO

MiMo-V2.5 310B MoE Released

·2026.04.29 03:27

Key point

XiaomiMiMo has released the non-Pro MiMo-V2.5 model with 310B/15B parameter scale.

1 / 2

Details

XiaomiMiMo has released MiMo-V2.5. It is a Sparse MoE-based omnimodal model, activating 15B parameters out of a total of 310B.

It processes text, images, video, and audio in a single architecture, with a context length of up to 1M tokens.

  • Compared to Pro, the non-Pro version lowers total parameters from 1.02T → 310B and active parameters from 42B → 15B.
  • The backbone consists of 48 layers (1 dense + 47 MoE), with 256 routed experts and 8 experts per token.
  • Training was conducted on approximately 48T tokens using FP8 mixed precision.
  • It is described as reducing KV cache storage by about 6x using an SWA:GA 5:1 ratio and a sliding window of 128.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.