Ask an Expert: How Does AI Understand Visual Search?
Key point
Google has enhanced multi-object visual search, enabling simultaneous search of multiple objects within an image using Gemini models.
Details
Existing visual search only allowed searching for one item at a time. However, with updates to Circle to Search and Lens, it is now possible to identify and search for multiple objects within an image simultaneously.
This feature is implemented through AI Mode, which is based on the multimodal capabilities of Gemini models. The AI analyzes the user's question together with the image to determine the best tool, finding multiple components within the image at the same time and integrating them into a single response.
The key is fan-out technology. After the AI model identifies multiple objects within an image, it runs multiple searches simultaneously, reads the results, and generates a coherent answer in just a few seconds.
Users can also start with a text search and then ask additional questions about a specific image, allowing for wide-ranging use from shopping to information exploration.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.