GUI Grounding Models Vista 9B/4B Released
·2026.06.13 16:50
Key point
Vista, a model based on Qwen 3.5 that precisely predicts the coordinates of GUI elements, has been released.
Details
inclusionAI has released VISTA-9B and VISTA-4B, GUI Grounding vision-language models using Qwen 3.5 9B as the backbone. These models take screenshots and natural language instructions as input and output click coordinates within the image frame as normalized values between 0-1000.
The key technical features are as follows:
- View-Consistent GRPO Training: By generating and training on different views of the same GUI instance, the model's ability to find consistent locations even in screenshots with geometric variations has been strengthened.
- Self-verified Cross-view Anchoring: A training method that adds an oracle-format center-point anchor only when the model's generated result achieves maximum reward, improving the stability of coordinate generation.
The models are provided as open source via Hugging Face.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.