AI Briefing
KO

Reasoning-Based Reward for Image Editing

·2026.05.04 09:00

Key point

Edit-R1 improves image editing performance with a reasoning-type reward model.

Details

Existing image editing reward models often produce only an overall score, failing to finely capture the different requirements of each instruction. Edit-R1 takes a verifier-based RRM approach, splitting instructions into multiple principles and verifying each item separately to create interpretable, fine-grained rewards.

The training pipeline consists of three stages.

  • Use SFT cold-start to first create a CoT reward trajectory.
  • Use GCPO to incorporate human pairwise preference data and strengthen the sample-level (pointwise) RRM.
  • Use the completed RRM as a non-differentiable reward to train the editing model with GRPO.

In experiments, Edit-RRM proved stronger as a dedicated editing reward model than Seed-1.5-VL and Seed-1.6-VL, and performance improved steadily as the model scaled from 3B to 7B. It also boosted the performance of editing models such as FLUX.1-kontext.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.