[CVPR 2026] Beyond Text: How VinQA Integrates Visual Elements into Multimodal Document QA Answers - LG AI Research BLOG
Key point
LG AI Research has unveiled VinQA, a study that generates answers by strategically combining text and visual elements.
Details
Existing multimodal document QA research has a limitation: even when the entire document is input or relevant pages are retrieved via RAG, the final answer is generated as text only.
To address this, LG AI Research proposed VinQA. Rather than simply attaching images at the end of an answer, VinQA aims to generate long-form answers with integrated visual elements, placing relevant diagrams or charts directly before the sentences that describe them.
Key features of VinQA are as follows:
- VinQA Dataset: A large-scale dataset spanning 7 domains, including academic papers, websites, and textbooks, was built. The training data consists of approximately 130,000 pages and 42,700 QAs.
- Multimodal RAG-based: Based on context retrieved through a multimodal RAG pipeline, the model generates answers combining text and visual citations.
- Two encoding methods: Model comprehension was improved through two approaches: Page Encoding, which preserves the layout of an entire page, and Modality Encoding, which extracts and processes visual elements.
This research was presented at CVPR 2026, and it offers a new direction on how visual materials can enhance the clarity of answers in real-world document environments.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.