Stable Virtual Camera: Multi-View Video Generation Using 3D Camera Control
Key point
Stable Virtual Camera is a general-purpose diffusion model that generates highly consistent multi-view videos through 3D camera control.
Details
Stable Virtual Camera is a general-purpose diffusion model that generates novel scenes based on multiple input views and target camera information. Unlike existing approaches that struggled with large viewpoint changes or maintaining temporal continuity, this model overcomes these issues through simple model design, an optimized training recipe, and a flexible sampling strategy.
Key features are as follows:
- Maintains high consistency without separate 3D representation-based distillation
- Capable of generating high-quality video up to 30 seconds long
- Supports seamless loop closure
Extensive benchmark results demonstrate that Stable Virtual Camera outperforms existing methodologies across various datasets and settings.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.