AI Briefing
KO

Nvidia unveils high-speed Grounding model

·2026.05.28 15:43

Key point

Nvidia has released LocateAnything, a Vision-Language Grounding model that is 10x faster than Qwen3-VL through parallel box decoding.

Details

Nvidia has released LocateAnything-3B, a Vision-Language Grounding model that precisely identifies the location of objects within images.

This model introduces Parallel Box Decoding technology, achieving up to 10x faster inference speed than the existing model Qwen3-VL while still delivering high-quality results.

Key features are as follows:

  • High-speed inference: Dramatically improved object localization speed through a parallel decoding approach.
  • High-quality Grounding: Accurately extracts object coordinates based on visual language understanding.
  • Accessibility: The model has been released via Hugging Face, and a demo can be tried immediately through GitHub and Spaces.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.