AI Briefing
KO

Qwen2.5-VL Released

·2025.01.26 20:08

Key point

Qwen has released Qwen2.5-VL, its next-generation vision-language model with agentic capabilities and long-video understanding.

Details

Qwen has released Qwen2.5-VL, its next-generation flagship vision-language model (VLM), marking a significant leap from its predecessor, Qwen2-VL. This release comes in three sizes—3B, 7B, and 72B—available as both Base and Instruct models.

Key features include:

  • Visual understanding: Proficient not only at recognizing common objects but also at analyzing text, charts, icons, and layouts.
  • Agentic capabilities: Functions as a visual agent capable of using computers and smartphones, dynamically directing tools and reasoning.
  • Long-video comprehension: Understands videos over 1 hour long and can accurately pinpoint specific events.
  • Precise localization and structured output: Accurately localizes objects using bounding boxes or points, and can convert data such as invoices or tables into structured data in JSON format.

In terms of performance, Qwen2.5-VL-72B-Instruct delivers competitive results against SOTA models across various benchmarks, including math, document understanding, and video understanding. It shows particular strength in document and chart comprehension, and can perform as a visual agent without additional fine-tuning.

For small and medium-sized models, Qwen2.5-VL-7B-Instruct outperforms GPT-4o-mini on multiple tasks. The 3B model, designed for edge AI, achieves better performance than the previous Qwen2-VL 7B model.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.