AI Briefing
KO

inclusionAI Releases Ling-3.0-flash-VL

·2026.09.09 00:57

Key point

inclusionAI has released Ling-3.0-flash-VL, a multimodal model featuring a 1M token context window and MoE architecture.

Details

inclusionAI has released Ling-3.0-flash-VL, a multimodal LLM that integrates understanding of text, images, and video. This model inherits the language and reasoning capabilities of Ling-3.0-flash while extending native image and video understanding features.

Architecture and Performance

The model holds a total of 124B parameters but activates only 5.5B parameters per token through a Mixture of Experts (MoE) structure to enhance inference efficiency. It supports a context window of up to 1M tokens and adopts a 42-layer hybrid backbone (alternating KDA and Gated MLA layers in a 5:1 ratio) to optimize long-context processing.

Vision Processing Technology

Visual features are aligned with text representations via a Vision Transformer (ViT) encoder and a 2-layer MLP projector. Notably, the application of VideoRoPE technology encodes spatial positions and temporal order, enabling tasks involving temporal changes such as event localization, long-video question answering, and video clip editing.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.