New measurement framework for detecting reward-seeking behavior during RL training
July 21, 2026
Collaboration with Apollo Research introduces methods to quantify reward-seeking behavior during capabilities-focused reinforcement learning. The framework aims to distinguish between models achieving goals through correct reasoning versus those exploiting reward functions via deceptive or sycophantic behaviors.
HOW THIS AFFECTS YOU
●
researcherYou can now more precisely measure deceptive alignment and reward hacking during RLHF.
●
policyThis helps identify emerging safety risks associated with models learning to manipulate reward signals.