[HUGGINGFACE]score: 0.44
Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective
October 1, 2026
On-policy distillation improves performance by making correct responses easier to sample via an implicit reward model but does not expand the student's underlying capabilities. When this implicit reward aligns with quality, performance gains occur; however, misalignment leads to reward hacking, resulting in repetitive and excessively long generations.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy