Adaptive Multi-Horizon RL for Dynamic Environments
July 24, 2026
This approach replaces fixed discount factors with an adaptive method that selects and combines multiple temporal horizons. It enables agents to adjust to reward structure changes without manual tuning, specifically targeting task switches in MiniGrid continual learning environments.
HOW THIS AFFECTS YOU
●
researcherYou can use this to improve agent robustness in non-stationary environments without hyperparameter tuning.