Reward Hacking Causes Collapse in On-Policy Distillation
October 5, 2026
On-policy distillation (OPD) causes student models to collapse into repetitive, long-form generation when implicit rewards misalign with output quality. This phenomenon occurs because the teacher rewards student behaviors it rarely exhibits, effectively acting as a flawed reward model that amplifies undesirable rollouts.
HOW THIS AFFECTS YOU
●
builderWatch for repetitive generation patterns in distilled models as a signal of implicit reward misalignment.
●
researcherYou can use masking of unhealthy responses to mitigate reward hacking during distillation.