AsymFlow, SOTA in pixel space
Key point
Stanford researchers presented AsymFlow, delivering more realistic AI images than latent diffusion.
Details
Stanford researchers proposed Asymmetric Flow Models (AsymFlow). It is a rank-asymmetric velocity parameterization for high-dimensional pixel-space generation, in which noise prediction is restricted to a low-rank subspace while data prediction retains full dimensionality.
This design analytically recovers the full velocity without changing the network architecture or the training/sampling procedure.
Key results are as follows.
- On ImageNet 256x256, it achieved FID 1.57.
- The authors stated it substantially outperformed existing DiT/JiT-family pixel diffusion models.
- They presented a path for fine-tuning a pretrained latent flow model into a pixel-space model.
- AsymFLUX.2 klein, based on FLUX.2 klein 9B, surpassed the latent base on HPSv3, DPG-Bench, and GenEval.
The authors explained that this approach preserves high-level semantics and structure while mainly correcting low-level details and textures.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.