AI Briefing
KO

The Reality of On-Policy Distillation: When It Helps, When It Hurts, and Why

·2026.07.09 09:00

Key point

This paper proposes a diagnostic framework that analyzes, at the token level, the impact of On-Policy Distillation on training reasoning models.

Details

On-Policy Distillation provides token-level dense supervision signals for training reasoning models, but it has not been clear under what conditions this signal is beneficial or harmful. Previously, confirming this required extremely costly training processes, and overall performance metrics alone made it difficult to grasp individual token-level dynamics.

This study introduces a Training-free diagnostic framework that operates per token, question, and Teacher model. By defining the parameter update that maximizes the student model's probability of success, the authors derive an Ideal per-node gradient, and develop a Scalable targeted-rollout algorithm to efficiently estimate it.

The analysis shows that the Distillation guidance exhibits a much higher Gradient alignment score for cases where the student model got the answer wrong (Incorrect rollouts) than for cases where it already got the answer right. This suggests that in correct-answer situations, the teacher's signal can actually act as noise.

Furthermore, the optimal Distillation context is jointly determined by the student model's capacity and the target task, meaning there is no single configuration applicable to all situations. Therefore, it is essential to perform task- and token-specific diagnostic analysis during the Distillation process.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.