AI Briefing
Sign in

Naver Cloud Releases Recipe for Training Image Generation Models Without Pre-training

·2026.09.29 09:00

Key point

Applying FLUX.2 VAE and Muon improves training efficiency and strengthens understanding of Korean instructions

1 / 11

Details

Naver Cloud has released the process of training large-scale image generation and editing models from scratch without pre-trained weights. It includes an experimental methodology and final recipe for designing and independently improving optimal model structures, data, and training objectives within limited resources.

Adoption of Key Techniques and Performance Improvements

The experiments were conducted by quickly validating with small models before applying them to large-scale training. FLUX.2 VAE, the Muon optimization algorithm, Self-Flow representation alignment, Long Skip, and VLM multi-layer text conditioning were adopted as the final configuration.

  • VAE Replacement: Replacing FLUX.1 VAE with FLUX.2 VAE improved FID by approximately 7% from 10.95 to 10.19 in ImageNet experiments, and reduced the time to reach FID 30 from 87.5k steps to 50k steps.
  • Optimization Algorithm: Applying Muon instead of AdamW adjusts the direction of matrix weight updates, significantly improving training efficiency. After 36 hours, FID improved from 10.19 with AdamW to 5.20 with Muon.
  • Representation Alignment and Structure: Applying Self-Flow and Long Skip improved DINO-MMD metrics in the later stages of training, and utilizing VLM multi-layer structures increased text rendering accuracy.

Noise Shift Values and Rejected Techniques

The Noise Shift values for the training and generation stages have different roles and must be optimized separately. Higher generation shift values tended to result in lower FID. On the other hand, JiT (x0-prediction), Register Token, and specific alignment method variations were not adopted due to negligible quality improvements or insufficient benefits relative to computational cost in the VAE latent space combination.

Service Application and Future Directions

The final recipe confirmed compliance with complex descriptive prompts, Korean and English text rendering, and image editability. Naver Cloud plans to apply this to its services, including the creation of business documents, presentation materials, and advertising assets. In particular, the core goal is to ensure accurate understanding of instructions, distinction between maintained and changed parts, and quality consistency during repeated requests, which are required in enterprise and public institution environments.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.