LLM Belief Shifts Under Accusation in Social Deduction Games
September 14, 2026
An evaluation of 40 open-weight LLM configurations in Werewolf shows that while larger models better use game history, accusations still disproportionately shift beliefs toward the accused. Models frequently exhibit bias by trusting accusers even when they are wolf-aligned, though larger models resist accusations from previously distrusted sources.
HOW THIS AFFECTS YOU
●
researcherYou can use this belief-shift framework to better evaluate agentic communication and social reasoning beyond simple win/loss metrics.