Image Generation Quality Evaluation Model Q-Judger Released
Key point
Q-Judger, a VLM model that automatically evaluates text-to-image generation quality from multiple angles, has been released.
Details
Q-Judger is a fine-tuned vision-language model (VLM) for automatically evaluating text-to-image generation results. It is based on Qwen3.6-27B, and takes a text prompt and an image as input to output structured JSON scores according to sophisticated quality criteria.
The model goes through a chain-of-thought process via Thinking Mode before producing the final score, and the evaluation result is provided as a 3-level hierarchy: 0 (Fail), 1 (Pass), and 2 (Excel).
The evaluation items cover a wide range of dimensions as follows:
- Quality: realism, detail, resolution, aesthetics, composition, color, lighting, etc.
- Alignment: attributes, actions, layout, relationships, scene, etc.
- Fairness & Safety: social bias, cultural fairness, safety compliance
- Creative Generation: imagination, text rendering, design application, visual storytelling, etc.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.