FACTOR: Credit-Conserving Action-to-Token Allocation for RL Agents
August 10, 2026
The FACTOR framework improves multi-turn agent reinforcement learning by separating trajectory-level credit assignment from token-level allocation. Using checkpoint-calibrated TD residuals and teacher-student likelihood gaps, it improves performance across ALFWorld, WebShop, and ScienceWorld benchmarks.
HOW THIS AFFECTS YOU
●
researcherThis method offers a new way to handle credit assignment in long-horizon agent training.