A new causal taxonomy distinguishes between deceptive behavior and underlying mechanisms like prior commitment and false preference. Testing on open-weight models shows that deceptive-looking behavior can occur without the proposed deceptive mechanisms.
HOW THIS AFFECTS YOU
●
researcherYou can use this framework to better differentiate between accidental behavior and intentional strategic deception.
●
policyThis framework helps define what constitutes 'deceptive' AI for safety and governance standards.