Model-Corrected World Models via Calibration-Risk Routing
October 2, 2026
The MC-WM framework mitigates simulator-to-target shifts in model-based reinforcement learning by partitioning target data into fit, selection, and calibration sets. It uses learned confidence signals and validity predicates to weight one-step imagined policy updates, reducing error when adapting to new dynamics.
HOW THIS AFFECTS YOU
●
researcherYou can use this method to stabilize model-based RL during domain shifts without rewriting physical reward functions.