RetireOPD: Adaptive Distillation for Agentic Reinforcement Learning
September 16, 2026
RetireOPD improves agent training by using an adaptive schedule to retire self-teacher supervision. It optimizes a skill-conditioned teacher first, then trains a skill-free student jointly with RL and on-policy distillation to ensure more reliable skill internalization.
HOW THIS AFFECTS YOU
●
builderThis recipe can help you train more robust agents that internalize complex skills without constant teacher dependency.
●
researcherThis addresses the stage-dependent benefits of teacher supervision in multi-turn agent RL training.