Recursive Self-Improvement via On-Policy Distillation for Reasoning
September 28, 2026
On-policy self-distillation (OPSD) enables recursive model improvement by using a frozen student model with ground-truth context as a teacher for a secondary student model. This method provides dense, token-level supervision during trajectory generation to improve reasoning capabilities without requiring an external teacher model.
HOW THIS AFFECTS YOU
●
researcherYou can use this framework to explore autonomous model scaling through self-generated, supervised trajectories.