AI Briefing
KO

High-Speed, High-Resolution Image Synthesis Using Latent Adversarial Diffusion Distillation

·2025.03.19 04:00

Key point

SD3-Turbo, which generates high-resolution images with just 4 sampling steps, was introduced through LADD technology.

Details

Diffusion models are a core technology for image and video generation, but they have the drawback of slow inference speed. Recently, the Adversarial Diffusion Distillation (ADD) method was introduced with the goal of single-step inference, but it had issues with difficult and costly optimization due to its reliance on a fixed DINOv2 discriminator.

The new method, Latent Adversarial Diffusion Distillation (LADD), overcomes these limitations. Unlike pixel-based ADD, LADD leverages the generative features of a pretrained Latent Diffusion Model. This approach simplifies the training process while boosting performance, enabling high-resolution and various aspect-ratio image synthesis.

The research team applied LADD to Stable Diffusion 3 (8B) to implement SD3-Turbo. SD3-Turbo achieves performance on par with state-of-the-art (SOTA) text-to-image generation models with just 4 steps of unguided sampling.

Additionally, the scaling behavior of LADD was systematically analyzed, and its effectiveness was demonstrated across various applications such as image editing and inpainting.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.