Unified Framework for On-Policy Self-Distillation via Learning Capacity Constraints
August 11, 2026
A new optimization framework for on-policy self-distillation (OPSD) couples token weighting with teacher-student divergence. By treating student learning capacity as a budget for aggregate learning difficulty, the method optimizes both the selection of tokens and the amount of privileged information internalized.
HOW THIS AFFECTS YOU
●
researcherYou can better balance token selection and information density during self-distillation tasks.