On-Policy Reverse Distillation for Weak-to-Strong Generalization
September 7, 2026
On-Policy Reverse Distillation (OPRD) enables stronger models to surpass weaker supervisors by amplifying verifier-driven policy gradients along the teacher's policy shift. This method avoids the capacity ceilings typically imposed by conventional distillation when treating a weak model as an optimization target.
HOW THIS AFFECTS YOU
●
researcherYou can train frontier models using weaker supervision without inheriting the teacher's performance limits.