AI Briefing
KO

STARFlow-V: End-to-End Video Generation Modeling Based on Normalizing Flows

·2026.04.30 09:00

Key point

STARFlow-V demonstrates high-quality autoregressive video generation without diffusion.

Details

STARFlow-V is an end-to-end model that applies normalizing flows (NF) to video generation. While naturally providing end-to-end training and likelihood estimation, it presents a new alternative to the diffusion-dominated video generation landscape.

Building on the recently proposed STARFlow, the model uses a global-local architecture that operates in a spatiotemporal latent space. It restricts causal dependencies to the global latent space alone while preserving local interactions within frames, reducing error accumulation over time.

On top of this, flow-score matching is added to improve the consistency of autoregressive generation with a lightweight causal denoiser, and video-aware Jacobi iteration reformulates internal updates into parallelizable iterations to boost sampling efficiency.

  • The same model supports text-to-video, image-to-video, and video-to-video.
  • In experiments, it showed strong visual quality and temporal consistency compared to diffusion-based baselines, along with practical sampling throughput.
  • The authors present this as the first evidence that normalizing flows are capable of high-quality autoregressive video generation, positioning it as a promising direction toward world models.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.