On-Policy Distillation Enables Weak-to-Strong Capability Transfer
September 25, 2026
Research shows that on-policy distillation (OPD) allows compact RL experts to transfer reasoning capabilities to much larger students. In weak-to-strong setups, students can even exceed the teacher's 'gold score' as performance rises linearly with the square root of reverse KL divergence.
HOW THIS AFFECTS YOU
●
researcherYou can leverage OPD to scale reasoning capabilities across different model architectures and sizes.
●
founderThis suggests an opportunity to use specialized, smaller models to train or enhance larger, more general models.