Inferring Visual Concepts from Image Sets
Key point
A model called VICIS has been proposed that enables Vision-Language Models to infer visual concepts from example images and apply them.
Details
Current Vision-Language Models (VLMs) follow complex text instructions well, but are weak at reasoning from purely visual information. In particular, they fail to grasp concepts shared across a set of images and apply them to new inputs.
To evaluate this ability, the research team defined the VICIS (Visual Concept Inference from Sets) task. Given a set of context images that share a specific concept and a query image, the model must generate a new image that matches the query while preserving the concept defined by the context.
Evaluation of SOTA VLMs showed poor performance. They tended to ignore the visual context or default to biased generation.
To address this, the research team proposed a training framework and architecture that infers visual concepts from image sets and extracts concept-specific embeddings. In experiments on synthetic data and large-scale ImageNet/WordNet data, the proposed model produced more accurate and diverse outputs, and was also confirmed to generalize to unseen concepts and new modalities such as sketches.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.