AI Briefing
KO

Xiaomi's MiMo Achieves 15x Faster Speed Than ChatGPT and Claude

·2026.06.09 09:00

Key point

Xiaomi has unveiled MiMo-V2.5-Pro-UltraSpeed, which generates 1,000 tokens per second by applying FP4 quantization and DFlash technology.

Details

Xiaomi and its inference partner TileRT developed the MiMo-V2.5-Pro-UltraSpeed model, which has 1 trillion parameters. This model recorded an overwhelming inference speed of 1,000 tokens per second on a standard 8-GPU commodity node environment.

This performance was achieved through the following technical innovations.

  • FP4 quantization: applied to the model's expert layers to maximize efficiency
  • DFlash speculative decoding: a method that proposes an entire block in a single pass instead of generating tokens one by one

The model offers a limited API trial from June 9 to June 23. The cost is 3 times higher than the standard MiMo-V2.5-Pro rate, but the output volume is about 10 times more.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.