Qwen3-Omni-Flash-2025-12-01: An Upgrade That Listens, Sees, and Follows More Intelligently
Key point
**Qwen3-Omni-Flash-2025-12-01** has been upgraded to handle speech, vision, and text more reliably across the board.
Details
Qwen3-Omni-Flash-2025-12-01 is an upgraded version of Qwen3-Omni, a multimodal model that jointly processes text, images, audio, and video. It generates text and natural speech simultaneously via real-time streaming, and this version raises both performance and efficiency.
The biggest change is the improvement in audio-visual interaction. It reduces the "intelligence degradation" problem that commonly appeared in everyday conversational scenarios, and improves the stability and consistency of multi-turn audio-visual conversations, enabling more naturally flowing interactions.
system prompt control has also been strengthened. System prompts can now be fully customized, allowing fine-grained control over details such as personality style (e.g., sweet, cool, anime-inspired), colloquial tone, and output length.
Language support has also expanded. Text-based interaction supports 119 languages, speech recognition supports 19 languages, and speech synthesis supports 10 languages, and the language compliance instability seen in the previous version is said to have been resolved.
Speech synthesis focuses on more human-like utterances. It reduces the problem of speech sounding slow or mechanical, and adjusts speaking rate, pause, and intonation according to text context to produce more natural and expressive speech.
Improvements have been reported across all modalities in terms of performance.
- Text understanding/generation: ZebraLogic +5.6, LiveCodeBench-v6 +9.3, MultiPL-E +2.7, WritingBench +2.2
- Speech understanding: reduced word error rate on Fleurs-zh, VoiceBench +3.2
- Image understanding: MMMU +4.7, MMMU-Pro +4.8, MathVision_full +2.2
- Video understanding: MLVU +1.6 and enhanced audio-visual synchronization
Going forward, the plan is to further develop multi-speaker ASR, video OCR, and audio-visual proactive learning, as well as expand support for agent-based workflows and function calling. The core message is "Hear You. See You. Follow Smarter.", aiming for more natural and precise multimodal interaction.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.