AI Briefing
KO

LG AI Research 455

·2026.07.16 09:00

Key point

It introduces the 'Generative Image Dynamics' research, selected as Best Paper at CVPR 2024, a technology that generates natural motion from still images.

1 / 2

Details

At CVPR 2024, a prestigious conference in the computer vision field, 2,719 papers were accepted, and among them, the image and video synthesis field received the hottest attention. In particular, this conference's Best Paper winner, 'Generative Image Dynamics', deals with technology that predicts natural motion from a single image to generate video.

This research focuses on implementing natural motion caused by external forces, such as trees, flowers, and candle flames swaying. Unlike existing studies that suffered from the problem of video consistency breaking down over time, this model performs analysis in the Frequency Domain to generate naturally repeating videos.

The key technical features are as follows:

  • Spectral Volume prediction: Uses a Diffusion Model to predict the motion of each pixel in the frequency domain.
  • Latent Diffusion Model based: Maps images into latent space via a VAE, and applies the denoising process to the Spectral Volume rather than RGB images.
  • Inverse Fourier Transform: Restores the predicted frequency-domain motion back into the target image.
  • Hierarchical structure: Uses a pyramid-shaped hierarchical structure to remove empty spaces that may occur during the warping process.

Furthermore, the practicality of the technology was demonstrated through an interactive demo where, when a user clicks and drags an image, motion occurs as if touching a real object and then returns to its original position.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.