How Cosmos 3 Helps Physical AI Make Decisions in Advance
Key point
NVIDIA has unveiled **Cosmos 3**, a world foundation model that supports the perception, prediction, and action of physical AI.
Details
For physical AI systems such as robots, autonomous vehicles (AVs), and smart spaces to operate autonomously, they must go beyond simply seeing the present and predict what will happen next. However, directly reproducing and training on complex real-world scenarios is extremely difficult in terms of cost and time.
NVIDIA Cosmos 3 is a new World Foundation Model designed to solve this problem. This model combines multimodal generation spanning text, video, image, ambient sound, and action, helping developers generate world data with physical context.
Based on a mixture-of-transformers architecture, Cosmos 3 operates as follows:
- Reasoning block: interprets what is currently happening in the scene.
- Generation block: generates physically grounded outputs, such as synthetic video or robot task data, based on the interpreted context.
In particular, Cosmos 3 is an omnimodel that can directly generate numerical data describing robot movement, such as joint angles, gripper positions, and trajectory points. This allows developers to fine-tune the model for specific robot morphologies or task environments.
Currently, the NVIDIA GEAR team is using this to develop video action models, and Agile Robots is leveraging Cosmos 3 to generate diverse task trajectory data for developing policies for humanoid robots. Cosmos 3 Nano has demonstrated outstanding performance, leading on the simulation-based RoboLab and the real-world RoboArena benchmarks.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.