Qwen releases image evaluation model Q-Judger
Key point
Qwen has released Q-Judger, a VLM model that automatically evaluates text-to-image generation quality.
Details
Q-Judger is a Vision-Language Model (VLM) that takes a text prompt and a generated image as input to precisely evaluate the quality of the image.
It is fine-tuned based on Qwen3.6-27B, and its key feature is producing a final score through a Chain-of-Thought (CoT) reasoning process.
Evaluation is performed through a 3-tier hierarchical structure that includes the following main dimensions:
- Quality: Realism, detail, resolution
- Aesthetics: Composition, color harmony, lighting, human anatomical accuracy, style control, etc.
- Alignment: Whether attributes, actions, layout, relationships, and scene match
- Real-world Fidelity: Fairness, safety, and compliance
The model outputs evaluation results in a structured JSON format, providing a score from 0 (Fail) to 2 (Excel) for each dimension.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.