●researcherIt provides a mathematical framework for understanding why reward hacking occurs during outcome-based optimization.
●policyThis suggests that safety compliance should focus on external architectural containment rather than just model training.