AI Briefing
KO

Microsoft unveils 4B image generation model Mage-Flow

·2026.07.23 03:05

Key point

Microsoft has unveiled Mage-Flow, an image generation and editing model that rivals large models with just 4B parameters.

Details

Microsoft has unveiled Mage-Flow, which handles text-to-image generation and instruction-based editing in a single stack. Despite its 4B scale, it shows competitive quality against much larger models like Qwen-Image 20B and FLUX.2 32B.

The core components are two-fold:

  • Mage-VAE: A lightweight tokenizer using 1-step diffusion encoding/decoding. Achieves ~12× reduction in encoding MACs and ~22× reduction in decoding MACs compared to FLUX.2-VAE
  • NR-MMDiT: A native-resolution multimodal DiT that supports a single checkpoint across 512-2048 resolutions, up to extreme 4:1 aspect ratios

With native-resolution packing and fused CUDA kernels, training speed has been improved from ~1.93s to ~0.78s (a roughly 2.5x reduction).

Three variants are provided—Base, RL-aligned, and 4-step Turbo—and both Mage-Flow for generation and Mage-Flow-Edit for editing are publicly available on Hugging Face.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.