Gemini Robotics-ER 1.6: Powering Real Robot Tasks with Enhanced Embodied Reasoning
Key point
Gemini Robotics-ER 1.6 strengthens spatial reasoning, instrument reading, and safety.
Details
Google DeepMind has released Gemini Robotics-ER 1.6. This model is a reasoning-first model designed so that robots go beyond simply following instructions, to understand and reason about the physical world.
The core improvements are spatial reasoning, multi-view understanding, task planning, and success detection. As a high-level reasoning model, it can even call Google Search, VLA(Vision-Language-Action) models, or user-defined external functions to carry out tasks.
In benchmarks, it showed overall better performance than Gemini Robotics-ER 1.5 and Gemini 3.0 Flash. In particular:
- pointing: more accurately points to multiple objects, and does not point to objects that are not present
- counting: improved quantity judgment and object distinction
- success detection: better determines task completion by jointly interpreting multiple camera viewpoints
As a new capability, it highlights instrument reading. This is a function for reading industrial-site equipment such as pressure gauges, sight glasses, and digital displays, reflecting real-world use cases from a collaboration with Boston Dynamics.
This capability doesn't just rely on simple visual recognition, but works through a process of:
- zooming into the image to check details
- estimating gauge markings and ratios via pointing and code execution
- applying world knowledge to interpret the final value
Safety has also been strengthened. The model has improved over the previous generation in compliance with Gemini safety policies, and has been trained to make safer choices in situations with physical constraints. It also outperformed the baseline, Gemini 3.0 Flash, on text/video safety evaluations based on real-life injury reports.
It is now available to developers via the Gemini API and Google AI Studio, along with a starter Colab example.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.