Beyond the 'Thinking Brain' to the 'Acting Brain'
Key point
We look at the two core technical directions of Robot Foundation Models (RFM), which go beyond LLMs in virtual worlds to interact with the physical environment.
Details
Recently, AI technology has been evolving beyond generating text and images into the era of Physical AI, which interacts with the real world. While existing LLMs were the 'thinking brain' of the virtual world, Physical AI refers to the 'acting brain' that perceives the physical environment and performs tasks through hardware.
At the center of this shift is the Robot Foundation Model (RFM), which can be commonly applied to various robots and tasks. While existing robot AI learned policies optimized for specific tasks, RFM aims to build general-purpose intelligence applicable to diverse environments and robot forms.
RFM research is largely unfolding in two technical directions.
- VLA (Vision-Language-Action) models: These directly connect visual and language information to robot actions. Representative examples include Gemini Robotics, NVIDIA GR00T N1, and Physical Intelligence π0, which are recently being advanced by incorporating reasoning and tactile signals, among other elements.
- World Model family: These predict future state changes after an action. They learn the dynamics of the physical world through video, complementing the physical change predictions that VLA is prone to miss.
Ultimately, the core of RFM lies in securing the ability to go beyond vision and language to predict changes in the actual physical world and generate sophisticated actions accordingly.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.