AI Briefing
KO

SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation

·2025.03.26 04:10

Key point

SV4D 2.0 is a Multi-View Video Diffusion model that generates high-quality dynamic 3D assets by enhancing spatio-temporal consistency.

Details

SV4D 2.0 is a Multi-View Video Diffusion model for dynamic 3D asset generation. Compared to its predecessor, SV4D, it is more robust to occlusion and large motion, and shows significantly improved generalization performance on real-world videos, detail sharpness, and spatio-temporal consistency.

Key improvements are as follows:

  • Network Architecture: Removed reference multi-view dependency and designed a blending mechanism for 3D and frame attention
  • Data: Enhanced both the quality and quantity of training data
  • Training Strategy: Adopted Progressive 3D-4D training to boost generalization performance
  • 4D Optimization: Addressed 3D inconsistency and large motion issues through 2-stage refinement and progressive frame sampling

Experimental results show that SV4D 2.0 achieved significant performance improvements both visually and quantitatively. In novel-view video synthesis, it achieved -14% LPIPS for detail and -44% FV4D for 4D consistency, and in 4D optimization it also recorded improvements of -12% LPIPS and -24% FV4D, respectively.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.