JiT-DDT Architecture Improves Text-Image Model Training Speed by 3.6x Compared to Linum v2
·2026.09.17 01:58
Key point
The JiT-DDT architecture, adopting a pixel-space encoder-decoder structure, achieved faster convergence and improved image quality while reducing GPU usage by 3.6x compared to Linum v2.
1 / 11
Details
The research team improved training efficiency by splitting the pixel-space DiT into an encoder-decoder structure (JiT-DDT). This architecture reduces GPU-hours by 3.6x compared to the existing Linum v2 while generating 4x more pixels, using an x-prediction method.
Key Technologies and Structure
- Encoder-Decoder Split: The encoder predicts the structure at 64x64 resolution, and the decoder generates 512x512 images using text conditions and the encoder's hidden states. The encoder and decoder are of the same size, and the decoder directly references text conditions.
- PixelREPA Application: Training stability was secured by aligning the encoder's hidden states using DINOv3 features.
Performance and Training Strategy
- Convergence and Quality: While the training speed is approximately 28% slower than JiT, the convergence speed is faster, and saturation phenomena are reduced, resulting in a final image quality that is visually assessed as 10-20% improved.
- Noise Schedule: The noise distribution was adjusted during the training process to induce detail recovery.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.