Steering Nash Equilibrium Selection via Reference Policy Anchoring
September 18, 2026
Anchoring a reference policy in regularized self-play allows for intentional steering toward specific equilibria in zero-sum games. Experiments on solvable games show mean coordinate errors as low as 0.007, proving that the reference policy, not initialization, dictates the final equilibrium selection.
HOW THIS AFFECTS YOU
●
researcherYou can use reference policy anchoring to control equilibrium selection in multi-agent training scenarios.