AI Briefing
KO

Qwen2.5 Omni: An All-in-One Multimodal Model That Sees, Hears, Speaks, and Writes

·2025.03.27 01:00

Key point

Qwen2.5-Omni, the new flagship model in the Qwen series, is an end-to-end multimodal model that unifiedly processes text, images, audio, and video.

1 / 2

Details

Qwen2.5-Omni, the new flagship model in the Qwen series, is an end-to-end multimodal model that processes diverse inputs including text, images, audio, and video in an integrated way. It provides simultaneous text generation and natural speech synthesis via real-time streaming.

The core is the Thinker-Talker architecture. Thinker acts as the 'brain', understanding inputs from various modalities and generating high-level representations and text, while Talker acts as the 'mouth', taking this output and producing natural speech in real time. In particular, to synchronize video input with audio timestamps, a new position embedding method called TMRoPE (Time-aligned Multimodal RoPE) was introduced.

Key features are as follows:

  • Real-time voice and video chat: Supports chunked input and immediate output, enabling fully real-time interaction.
  • Strong multimodal performance: Surpasses Qwen2-Audio in audio capability, and shows performance on par with Qwen2.5-VL-7B.
  • Excellent voice instruction following: Demonstrated voice instruction-following ability on par with text input on benchmarks such as MMLU and GSM8K.

It achieved state-of-the-art (SOTA) performance on OmniBench, and also recorded outstanding results on various single-modality tasks such as speech recognition (Common Voice), image reasoning (MMMU), and video understanding (MVBench).

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.