Testing six LLMs against 391 human participants shows that models fail to simulate initial stances or produce faithful belief updates unless provided with the actual human starting points. While Qwen3-32B and GPT-5-Mini can match post-discussion distributions, all models exhibit systematic biases toward neutral positions.
HOW THIS AFFECTS YOU
●
researcherAvoid using LLMs as proxies for human subjects in social science experiments without providing ground-truth initial states.