AI Briefing
KO

Qwen3.5-Omni: Scaling Up Toward Native Omni-Modal AGI

·2026.03.30 05:00

Key point

Qwen3.5-Omni significantly boosts omnimodal performance with a 256k context, 10-hour audio input, and 113-language ASR.

Details

Qwen3.5-Omni is the latest fully omnimodal LLM that handles text, image, audio, and audio-visual content all together, offered in three Instruct versions: Plus / Flash / Light. It supports a 256k long-context, and can process over 10 hours of audio input and over 400 seconds of 720P audio-visual input (1 FPS).

Pretraining was conducted in a native omnimodal manner, incorporating over 100 million hours of audio-visual data in addition to text and visual data. As a result, cognition and generation capabilities were strengthened across all modalities, and in particular, multilingual capability improved significantly compared to the previous generation, Qwen3-Omni, now supporting 113 languages/dialects for ASR and 36 languages/dialects for speech generation.

In terms of performance, Qwen3.5-Omni-Plus is reported to achieve SOTA across 215 benchmarks/subtasks covering audio and audio-visual understanding, reasoning, and interaction. It surpasses Gemini-3.1 Pro across overall audio understanding, has risen to the same level in audio-visual understanding, and matches the visual and textual capabilities of the same-sized Qwen3.5 series.

Notably standout features include the following.

  • Audio/Audio-Visual Captioning: Generates structured, detailed captions including automatic segmentation, timestamps, and even character relationships
  • Audio-Visual Vibe Coding: A new capability that performs coding directly from audio-visual instructions alone
  • Realtime API: Supports semantic-based interruption, WebSearch, complex FunctionCall, end-to-end voice control, and voice cloning
  • ARIA: Reduces omissions, misreadings, and unstable number pronunciation in streaming speech by dynamically aligning text and speech units

The architecture retains the existing Thinker-Talker structure, but the Thinker receives and processes visual and speech signals via a Vision Encoder and AuT, while the Talker performs context-based speech generation based on the Thinker's multimodal inputs and text outputs. For speech representation, RVQ is used instead of a heavy DiT, and chunk-wise streaming input combined with a streaming Talker design enables real-time interaction.

Compared to the previous generation, the Backbone changed from MoE to Hybrid-MoE, and the sequence length was expanded from 32k to 256k. Speech recognition coverage has also expanded significantly, securing much broader coverage including numerous languages and Chinese dialects.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.