Regularized Emphatic TD Learning for Stable Off-Policy Updates
September 18, 2026
RETD introduces a normalized first-order post-shock repair to stabilize emphatic temporal-difference learning under constant stepsizes. The method uses a leaky scalar state to provide delayed corrections, ensuring convergence where standard ETD may fail due to positive Lyapunov exponents.
HOW THIS AFFECTS YOU
●
researcherThis offers a mathematical solution for stability in off-policy TD learning with constant stepsizes.