AI Briefing
KO

Multimodal Reasoning Benchmark 'ConTextual' Released

·2024.03.05 09:00

Key point

A new dataset and leaderboard called 'ConTextual' has been released to evaluate multimodal models' ability to jointly reason over text and images.

1 / 2

Details

ConTextual, developed by UCLA researchers, is a dataset designed to evaluate the ability to reason by combining visual context and textual information within images that contain text.

While existing LMM (Large Multimodal Models) evaluations have focused on simple instruction following, ConTextual addresses context understanding in real-world text-rich scenarios such as map navigation, shopping, and infographics. The dataset has the following characteristics.

  • 8 real-world scenarios: Includes time reading, shopping, navigation, mobile apps, webpages, infographics, and more.
  • Dataset composition: Provides 506 challenging instructions along with a validation set (100) and a test set.

In experiments, the researchers tested 13 models, including closed models such as GPT-4V and Gemini-Vision-Pro, as well as open-source models like LLaVA and Qwen-VL. Evaluation uses an LLM-as-a-judge approach powered by GPT-4, and the researchers also released a leaderboard where model-specific performance can be viewed.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.