AI Briefing
KO

Ling-2.6-flash Open-Sourced

·2026.04.29 00:43

Key point

Ling-2.6-flash has been open-sourced, achieving up to 340 tok/s on 4× H20.

1 / 2

Details

Ling-2.6-flash has been officially open-sourced. It is an instruct model with 104B total parameters and 7.4B active parameters.

  • It replaces the existing GQA with a 1:7 MLA + Lightning Linear hybrid architecture and combines it with sparse MoE to boost inference efficiency.
  • On a 4× H20-3e setup (TP=4, Batch Size = 32), it recorded up to 340 tok/s, and prefill and decode throughput were reported to improve by up to about 4x.
  • In the overall Artificial Analysis evaluation, it reportedly maintained competitive performance while using only 15M tokens.
  • It reportedly showed strengths in agent benchmarks such as BFCL-V4, TAU2-bench, SWE-bench Verified, Claw-Eval, PinchBench, as well as in general knowledge, math reasoning, instruction following, and long-context understanding.
  • It provides a recommended SGLang configuration along with vLLM support, and for the MTP (Multi-Token Prediction) version, a separate patch branch was noted due to a bug in official SGLang.
  • Remaining limitations noted include tool hallucination in complex scenarios, Chinese-English switching, and the need for improvement in following complex instructions.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.