World Labs Unveils 'Atlas,' a Multimodal Model for 3D Spatial Intelligence
Key point
World Labs has unveiled Atlas, a next-generation world model that integrates text, image, video, and 3D data.
Details
World Labs has unveiled Atlas, a multimodal autoregressive diffusion transformer model that natively processes text, image, video, and 3D data. The model is designed to combine all inputs into a shared spatial context to maintain 3D consistency while generating the next frame.
Key Features and Capabilities
- Camera-Controlled Generation: Generates new views from user-specified camera positions and angles based on reference images, naturally extending into unobserved areas.
- Pixel-Level Camera Control: Uses precise camera geometry as native input instead of coarse text instructions, allowing fine-grained control over shot composition and motion.
- Spatial Context-Based Generation: Forms context by anchoring each image to a specific location in 3D space. When two unrelated images are placed in 3D space, Atlas generates a world that smoothly connects the space between them.
- Spatial Reconstruction: Can faithfully reconstruct real-world spaces from as few as 2–3 images, and supports up to 100+ images for precise environment restoration. This is possible without specialized equipment or hundreds of dense views.
Atlas is slated for integration into World Labs' upcoming products, including Marble, and features scalability that improves performance as training compute resources increase.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.