AI Briefing
KO

Introducing Qwen-VL

·2024.01.25 14:33

Key point

Qwen has released Qwen-VL-Plus and Qwen-VL-Max, multimodal models with greatly enhanced image reasoning and recognition capabilities.

Details

The Qwen-VL series, which expands Qwen's multimodal capabilities, has been upgraded into two enhanced versions: Qwen-VL-Plus and Qwen-VL-Max. This update significantly improves image-related reasoning capabilities as well as the ability to recognize, extract, and analyze text and detailed information within images.

In particular, it supports high-resolution images over 1 million pixels and various aspect ratios.

  • qwen-vl-plus: A model with enhanced detail recognition and text recognition capabilities, supporting ultra-high-resolution images and arbitrary aspect ratios.
  • qwen-vl-max: The most powerful model, with further improved visual reasoning and instruction-following capabilities, delivering optimal performance on complex tasks.

In terms of performance, these models match Gemini Ultra and GPT-4V across multiple multimodal tasks, significantly surpassing the records of existing open-source models. In particular, on Chinese question-answering and text comprehension tasks, they outperform GPT-4V and Gemini.

In real-world use, they have also demonstrated outstanding problem-solving abilities in conversation, identifying celebrities and landmarks, text generation, and interpreting visual content.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.