Ling-3.0-Flash Released
·2026.08.05 12:30
Key point
InclusionAI released a 124B MoE model with 5.1B active parameters.
Details
InclusionAI, under Ant Group, released the Ling-3.0-Flash weights under the MIT License.
- Total parameters are 124B, with approximately 5.1B parameters activated per token.
- It activates 8 out of 512 routing experts and 1 shared expert.
- It uses native hybrid linear attention, interleaving 35 KDA layers and 7 gated MLA layers in a 5:1 ratio.
- The research team claims benchmark performance is comparable to or exceeds their own 1T parameter Ring-2.6-1T. However, all figures are self-reported by the development team.
At launch, GGUF and llama.cpp are not supported, and official serving examples require their own SGLang and vLLM forks and 4 GPUs. Although the active parameters are small, the full 124B weights must be held in memory, creating a distinction between inference costs and model hosting costs.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.