Qwen Releases Qwen3.8-Omni-Flash with Enhanced Real-Time Audio-Visual Interaction and Agent Capabilities
Key point
Qwen has released Qwen3.8-Omni-Flash, featuring enhanced real-time audio-visual interaction and agent capabilities, along with related open-source tools.
Details
The Qwen Team has launched Qwen3.8-Omni-Flash for deep understanding and generation of full audio-visual content, including text, images, audio, and video. The model additionally introduces Qwen3.8-Omni-Flash-Realtime for low-latency real-time interaction, enabling it to call tools and perform tasks while receiving live audio-visual streams. Notably, it supports modeling that combines pronunciation and meaning through speech recognition and generation, as well as locating objects solely by sound via spatial audio recognition.
Agent Performance and Benchmarks
According to benchmark results, the model scored 71.0 on WildClawBench-MM, a multimodal tool-use evaluation, and 36.8 on AgenticVBench, showing performance improvements over the previous model (Qwen3.5-Omni-Plus). It also achieved a score of 74.0 on OmniGAIA, an evaluation of web search capabilities.
Expanding the Open-Source Ecosystem
Qwen released Omni Skill Creator as a new open-source feature of Qwen-MM-Plugins for agent skill generation. This allows users to extract Standard Operating Procedures (SOPs) from demonstration videos or learn expert tool usage and decision-making criteria to create reusable agent skills.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.