AI Briefing
KO

HY-World 2.0: Multimodal 3D World Model

·2026.04.20 09:00

Key point

HY-World 2.0 creates editable 3D worlds from text, images, and video.

1 / 2

Details

HY-World 2.0 is a multimodal world model that takes text, a single image, multi-view images, or video as input and generates/reconstructs a 3D world representation. The output is not simply video, but actual 3D assets such as mesh and Gaussian Splattings (3DGS), so it can be directly imported into Blender, Unity, Unreal Engine, and Isaac Sim.

There are two core axes.

  • World Generation: Given text or a single image, it creates a navigable 3D world through a four-stage pipeline.
  • World Reconstruction: Given multi-view images or video, it jointly predicts depth, surface normal, camera parameters, 3D point cloud, and 3DGS attributes in a single forward pass.

The generation pipeline consists of HY-Pano 2.0's panorama generation, WorldNav's trajectory planning, WorldStereo 2.0's world expansion, and world composition via WorldMirror 2.0 and 3DGS learning. Through this structure, it produces high-quality, navigable 3D scenes and expands them into worlds that can actually be moved through and explored.

The project strongly emphasizes its difference from existing video world models. Video models ultimately only produce "video to be played back," whereas HY-World 2.0 creates editable, persistent 3D assets. This gives it more direct practical utility in terms of 3D consistency, real-time rendering, physics collision, lighting, and engine compatibility.

WorldMirror 2.0 is particularly practical. From multi-view images or casual video, it predicts dense point cloud, depth map, surface normal, camera parameters, and 3DGS all at once, and it states that it supports flexible-resolution inference from 50K to 500K pixels. In other words, it is focused on quickly reconstructing digital twins from photos or video.

According to the news item, the technical report and some code were released on April 16, 2026, and on the same day, WorldMirror 2.0 inference code and model weights were also open-sourced. Meanwhile, the full World Generation inference code, HY-Pano 2.0, WorldNav, and WorldStereo 2.0 remain listed as coming soon.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.