Mathematical Conditions for Catastrophic Value Misalignment
August 3, 2026
This paper identifies specific conditions under which an agent with an $\eta$-catastrophic value function will be deployed despite idealized alignment training. It demonstrates that optimizing for imperfect proxies can lead to outcomes where expected human value drops below a defined threshold.
HOW THIS AFFECTS YOU
●
researcherYou should consider limiting optimization pressure through methods like quantilization to avoid proxy-driven value collapse.
●
policyThis provides a formal framework for understanding the risks of overoptimization in automated governance.