NVIDIA Unveils 'Cosmos 3,' an Omni Model for Physical AI
Key point
NVIDIA has unveiled Cosmos 3, a World Foundation Model (WFM) that unifies reasoning and action for physical AI.
Details
NVIDIA has released Cosmos 3, a World Foundation Model (WFM) for physical AI, on Hugging Face. Cosmos 3 is the first Omni-model to unify world generation, physical reasoning, and action generation into a single model.
Previously, separate models had to be combined for world generation, scene understanding, policy generation, and so on, but Cosmos 3 processes all modalities—text, image, video, audio, and action—within a single model through a Mixture-of-Transformers (MoT) architecture.
Core Technology and Composition:
- Unified Architecture: The AR (Autoregressive) subsequence, responsible for reasoning, and the DM (Diffusion) subsequence, responsible for generation, interact through Joint Attention, allowing flexible switching into a VLM, video generator, or robot policy model.
- Model Lineup: Cosmos 3 Nano (8B parameters), optimized for efficient inference, and Cosmos 3 Super, a high-performance model, are provided.
- Development Ecosystem: Along with model distribution via Hugging Face, post-training scripts for training on custom data and open synthetic data generation (SDG) datasets are supported on GitHub.
The model is expected to serve as a key foundation for building AI systems that need to understand and interact with the physical world, such as robotics, autonomous driving, and smart space simulation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.