AI Briefing
KO

[CVPR 2026] Beyond Text: VinQA, a Multimodal Document QA That Answers Using Visual Elements Like Diagrams, Charts, and Tables

·2026.07.20 17:00

Key point

LG AI Research has unveiled 'VinQA,' a multimodal document QA study that generates long-form answers by appropriately placing text and visual elements together.

1 / 2

Details

Existing multimodal LLM-based document QA has had the limitation of generating answers only as text. However, real documents contain a mix of various visual elements such as diagrams, charts, and tables, and appropriately utilizing them is key to improving the comprehensibility of answers.

VinQA, proposed by LG AI Research, is a study that goes beyond simply attaching images, generating Long-Form Answers by placing supporting visual elements at appropriate positions. To this end, the research team achieved the following results:

  • Construction of the VinQA dataset: The team reconstructed approximately 130,000 pages of data from 7 domains, including academic papers and guidebooks, to build a long-form answer dataset combining text and visual elements.
  • Study of two encoding methods: The team conducted a comparative study of Page Encoding, which encodes the entire page, and Modality Encoding, which uses OCR and cropped images.
  • New evaluation framework: The team proposed evaluation metrics such as M-GroSE and Visual G-Eval, which can measure the citation accuracy and placement appropriateness of visual elements.

Experimental results showed that when the open-source model Qwen2.5-VL-7B was fine-tuned on the VinQA dataset, it significantly narrowed the performance gap with commercial models GPT-4.1 and Claude 3.5 Sonnet.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.