AI Briefing
KO

Moonshot AI Releases Open Weights for 2.8T-Scale Kimi K3

·2026.07.29 09:30

Key point

Moonshot AI has released the weights and technical documentation for Kimi K3, an MoE model with 2.78 trillion parameters.

Details

Moonshot AI has released the weights and a 47-page technical document via Hugging Face for Kimi K3, a native multimodal MoE (Mixture of Experts) model with a total of 2.78 trillion parameters (104.2 billion active).

Beyond the scale race among open models, Kimi K3 focuses on raising the pre-training foundation to the 3T class while also emphasizing reinforcement learning (RL), reasoning, and support for a 1 million token long context. Notably, it achieved approximately 2.5x improvement in scaling efficiency compared to Kimi K2.

The key architectural features are as follows:

  • Three expansions of information flow: expands information across sequence length (KDA), network depth (AttnRes), and model width (Stable LatentMoE).
  • Hybrid attention: mixes KDA (69) and Gated MLA (24) layers.
  • Stable LatentMoE: adopts a structure that increases computational efficiency while expanding the model's width.
  • Native multimodal: integrates the MoonViT-V2 vision encoder to process visual information.
  • Quantization technology: applies quantization-aware training (QAT) using MXFP4 weights and MXFP8 activations.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.