Stochastic Teacher Intervention Reduces Error Accumulation in Agentic Distillation
October 9, 2026
STI-OPD introduces a framework for multi-turn agentic on-policy distillation that uses teacher intervention to prevent student error drift. By replacing student actions with teacher actions based on policy discrepancy, the method ensures the student receives reliable supervision during long-horizon tasks.
HOW THIS AFFECTS YOU
●
builderYou can use this approach to create more robust student models that maintain performance over long sequences of actions.
●
researcherThis offers a new method to stabilize on-policy distillation for complex, multi-turn agentic rollouts.