MiniMax previews M3 model with new sparse attention mechanism and 15.6x faster long-context response speed
Key point
MiniMax has previewed its M3 model, which introduces a new sparse attention mechanism to achieve 15.6x faster long-context response speed.
Details
China's AI company MiniMax has released a technical report on its popular M2 series model, previewing the key technology behind its next-generation M3 model. The M2 series employs a MoE (Mixture-of-Experts) architecture, holding 229.9 billion parameters while activating only 9.8 billion parameters per token, achieving efficient operation.
The M2 model adopts full multi-head attention with GQA (Grouped Query Attention) applied across all 62 layers. This enables precise context understanding, but causes a 'quadratic scaling' problem where computation increases exponentially as input length grows.
To address this, MiniMax is introducing a new sparse attention approach in its next-generation M3 model. Adopting a custom sub-quadratic framework, M3 aims to achieve the following performance improvements:
- 15.6x faster decoding speed: Dramatically improves response speed in long-context environments of 1 million tokens.
- Economical AI agent deployment: Lowers the hardware costs required for processing ultra-long context, securing economic viability for AI agent utilization.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.