Neurally Calibrated Reward Shaping for Prosocial Multi-Agent RL
August 6, 2026
Researchers recovered a guilt weight of 1.118 from human fMRI data to calibrate reward shaping in multi-agent reinforcement learning. Using PPO actor-critics in a Social Lottery environment, neurally calibrated agents most closely matched human social safe-choice rates.
HOW THIS AFFECTS YOU
●
researcherYou can leverage human neural data to move beyond hand-tuned social reward coefficients in MARL.