Inference Engine Support Status for Ling-3.0-flash
·2026.07.28 01:33
Key point
This summarizes the support plans and technical bottlenecks for the Ling-3.0-flash model across major inference engines.
Details
The support status of major inference engines for the recently released Ling-3.0-flash model is as follows.
- SGLang: Promised Day-0 support, and is developing the model based on its identified KDA + MLA hybrid attention structure.
- vLLM: Stated that support will be provided once the model weights are open-sourced.
- llama.cpp: Currently experiencing the biggest bottleneck. The support request for the Bailing MoE variant structure adopted by this model was closed as 'not_planned'.
Technically, the delta-net/KDA kernel is already implemented, so the attention structure itself is not an issue, but a PR to implement the Bailing MoE conversion path is needed.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.