Trajectory Intervention Self-Distillation for Enhanced On-Policy Learning
September 28, 2026
TISD addresses the data-collection bottleneck in on-policy self-distillation by using teacher-student disagreement to propose trajectory branches rather than just local repairs. The algorithm forces teacher-selected actions and regenerates successor contexts to provide denser supervision for the student model.
HOW THIS AFFECTS YOU
●
researcherYou can improve distillation efficiency by intervening on entire trajectories rather than just correcting single tokens.