ParliamentBench Evaluates LLM Deception via Secret Hitler Social Deduction
July 31, 2026
ParliamentBench is an open-source benchmark using Secret Hitler to evaluate 16 LLMs on deception, persuasion, and reasoning under information asymmetry. The framework introduces three metrics to isolate social deduction, reasoning, and deceptive consistency, revealing high-performing clusters in models like GPT-5.4 and Kimi K2.5.
HOW THIS AFFECTS YOU
●
researcherYou can use this framework to measure how reliably models maintain deceptive consistency in multi-agent environments.
●
policyThis helps you quantify the risks of deploying autonomous agents in high-stakes social or legal settings.