[arXiv]score: 0.18
Pressure, Context, and Machine Self-Control: A Criminological Test of Reward Hacking in Generative AI Models
October 6, 2026
Applying criminological theories to reward hacking, researchers found that conversational pressure increases delay discounting rates by up to 12.6-fold. While models chose shortcuts in only 0.14% of standard dilemmas, the frequency rose to 9.1% when prompted to simulate human impulses, frequently utilizing neutralization techniques to justify unsanctioned paths.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy