Video to 3D
Key point
HY-World 2.0 has been released, creating editable 3D worlds from text, images, and video.
Details
HY-World 2.0 is a multimodal world model that generates and reconstructs 3D worlds from text, single-image, multi-view image, and video inputs.
The core point is that instead of simple video generation, it directly creates editable 3D assets such as mesh and 3DGS (Gaussian Splatting).
- World Generation: Constructs 3D worlds through a 4-stage pipeline: text/image → panorama generation → path planning → world expansion → world composition
- World Reconstruction: From multi-view images/video, simultaneously predicts depth, surface normal, camera parameter, 3D point cloud, and 3DGS attribute in a single forward pass
- Practical orientation: Aims to produce forms that can be directly imported into engines such as Blender, Unity, Unreal Engine, and Isaac Sim
In this release, the WorldMirror 2.0 inference code and model weights have been open-sourced, along with a technical report. However, the full World Generation inference code, HY-Pano 2.0, and WorldStereo 2.0 related code are planned for future release.
The official description claims that HY-World 2.0 is comparable to closed-source SOTA-level 3D world models, and specifically emphasizes strengths in persistence, 3D consistency, real-time rendering, and physics-based interaction compared to existing video world models.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.