Qwen3.5-Omni Technical Report
Key point
Qwen3.5-Omni achieves SOTA with 256k context and over 100 million hours of audio-visual data.
Details
Qwen3.5-Omni is the Qwen Team's latest fully omnimodal LLM, capable of jointly understanding and generating text, images, audio, and audio-visual content. It supports hundreds of billions of parameters in scale and 256k context, and was pretrained on heterogeneous text-vision pairs and over 100 million hours of audio-visual data.
The model lineup is split into Plus and Flash, both instruct models. The report states that Qwen3.5-Omni-Plus achieved SOTA across 215 audio and audio-visual understanding, reasoning, and interaction subtasks and benchmarks, explaining that it surpassed Gemini-3.1 Pro on key audio tasks and reached its level on comprehensive audio-visual understanding.
The architecture retains the existing Thinker-Talker design while applying Hybrid-Attention MoE to both sides. This design improves long-sequence reasoning efficiency, enabling understanding of audio over 10 hours long and processing of 720P video (1 FPS) up to 400 seconds long.
On the speech generation side, it supports single-frame instant synthesis using multi-codebook codec representation. To reduce instability and unnaturalness in streaming speech synthesis, ARIA (Adaptive Rate Interleave Alignment) was incorporated, dynamically aligning the speed difference between text and speech units to improve natural prosody and stability while minimizing added latency.
Multilingual coverage is also extensive. Speech generation operates with emotional nuance across 10 languages, and training coverage extends to ASR in 113 languages/dialects and speech synthesis in 36. Zero-shot voice customization is also possible, matching a voice from just a user sample.
Key new features include the following.
- Controllable audio-visual captioning: structured captions including automatic segmentation, timestamps, and even character relationships, down to screenplay-level descriptions
- Real-time interaction: semantic interruption, control over volume, speed, and emotion, and voice cloning
- Native omnimodal agent behavior: WebSearch, complex FunctionCall, and Audio-Visual Vibe Coding, which generates executable code from audio-visual instructions alone
Above all, it's significant that text and vision performance did not degrade compared to unimodal Qwen-series models of the same size. It is offered via a public API, and Qwen3.5-Omni clearly shows a direction of combining audio SOTA, real-time conversation, and agentic behavior into a single model.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.