Qwen-Image-2.0 Technical Report
Key point
Qwen-Image-2.0 presents an image foundation model that unifies generation and editing.
Details
Qwen-Image-2.0 is an image generation foundation model that unifies generation and editing into one. It simultaneously targets ultra-long-text rendering, multilingual typography, high-resolution photorealism, complex instruction following, and efficient deployment—areas where existing models were weak.
The core architecture uses Qwen3-VL as the condition encoder and jointly models condition-target with a Multimodal Diffusion Transformer. This is combined with large-scale data curation and a customized multi-stage training pipeline to boost both understanding and generation/editing flexibility.
It can take instructions of up to 1K tokens to generate text-heavy content such as slides, posters, infographics, and comics. Multilingual text fidelity and typography have been improved, and textures, lighting, and details are rendered more realistically.
- It follows complex prompts more reliably.
- In human evaluations, it outperformed the previous Qwen-Image model in both generation and editing.
- It points toward a more general and reliable, practical model.