RFM: Action-Centric Intelligence for Physical AI - LG AI Research Blog
Key point
This post covers the concept and key technical directions of Robot Foundation Models (RFM), robots that perform tasks by directly interacting within physical environments.
Details
While existing AI has been a 'thinking brain' that generates text and images, Physical AI aims to be an 'acting brain' that perceives environments and performs tasks through real hardware. At the center of this shift is RFM (Robot Foundation Model), a general-purpose robotic intelligence applicable to diverse robots and environments.
Existing robot learning methods had the limitation of overfitting to specific robots or environments, causing performance to drop sharply with even small changes. RFM, on the other hand, must go beyond simply scaling language/vision models and possess the ability to interact with the physical world, observe the results, and dynamically modify its actions.
Current RFM research is unfolding along two major technical paths.
- VLA (Vision-Language-Action) models: This approach directly maps visual and language inputs to robot actions. Recently, beyond simple command execution, these models have been refining control policies by integrating reasoning, task context, and tactile/force signals.
- World Model (or World Action Model) approach: This focuses on predicting future environmental changes simultaneously with action generation. It plays a role in compensating for the limitations VLA models have in predicting physical and dynamic environmental changes.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.