AI Briefing
KO

Qwen Releases 'Qwen-Image-2.1', a 7B Parameter Unified Image Model

·2026.09.20 21:00

Key point

Qwen has released Qwen-Image-2.1, an open-source image model with 7 billion parameters that integrates generation and editing.

1 / 11

Details

The Qwen team has released Qwen-Image-2.1, an open-source model that unifies text-to-image generation and image editing capabilities. It features a compact architecture with only 7B parameters in the visual generation component, balancing generation quality, inference efficiency, and cost.

Efficient Architecture and Transparency Support

The model consists of 32 Single-Stream DiT layers and applies mixed-granularity attention to increase inference speed. It uses token-level causal masks for text and chunk-level masks for image generation, reusing the KV cache to reduce memory usage. Additionally, it integrates features from the previous dedicated model (Qwen-Image-Layered) to natively support the generation and editing of images with transparent backgrounds (RGBA) based on prompts.

Flexible Editing Features

It can combine up to 10 reference images to generate group photos, outfit combinations, interior layouts, and more. For Local Editing, it allows precise modification of specific areas using circles, paint annotations, or separate masks, showing strength in maintaining Fidelity for facial features in portraits or product textures. It is also applicable to various visual storytelling tasks such as panoramas, infographics, and storyboards.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.