AgentOPSD: Recursive Self-Distillation for Agentic RL
August 5, 2026
AgentOPSD introduces a critic-free, recursive method for turn-level credit assignment in long-horizon agentic tasks. It uses a Bayesian belief state in log-odds space to aggregate token-level teacher-student log-probability gaps into turn-level evidence.
HOW THIS AFFECTS YOU
●
researcherThis provides a new approach to solving the sparse reward problem in multi-turn agent training.