AI Briefing
KOSign in

UniEvo-VL: Multimodal Models Improve via On-policy Self-Distillation

·2026.10.05 09:00

Key point

The framework improved GenEval scores from 0.747 to 0.808 on Qwen-image-2512 by using the model's own critiques as training signals.

Details

Modern multimodal models can generate and understand content within a single system, allowing them to provide and learn from their own feedback. UniEvo-VL leverages this capacity by introducing a self-evolving framework where the model acts as both teacher and student, eliminating the need for a separate, larger external teacher.

Self-Distillation Mechanism

The framework uses the model's self-critiques as privileged information during training. The student model sees only the vanilla question, while the teacher model conditions on the privileged critique. Training minimizes the per-state divergence between their denoising diffusion distributions over the student's own sampling trajectories.

Performance Gains

Experiments conducted on the open-source Qwen-image-2512 demonstrated significant improvements in image generation capabilities:

  • GenEval: Score increased from 0.747 to 0.808.
  • GenEval2 Soft-TIFA: Score increased from 32.97 to 35.53.

Limitations and Observations

While the method improves overall generation, results indicate that self-improvement is not uniform across all tasks, particularly in mixed text-rendering outcomes. Additionally, tests with more powerful external critics like GPT5.6-Luna suggest that models with stronger judge capabilities can anticipate a higher ceiling for self-evolution.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.