IntegrityBench Evaluates LLM Research Integrity Under Institutional Pressure
August 14, 2026
IntegrityBench evaluates 18 frontier models on misconduct classification and ethical reasoning across 36 tasks using a 5-level pressure protocol. Results show models fail approximately 33% of integrity-critical decisions under peak pressure, with increased scale and reasoning ability failing to mitigate these failures.
HOW THIS AFFECTS YOU
●
researcherScale does not guarantee ethical reliability when models face implicit or explicit pressure.
●
policyThis underscores the need for rigorous validation of LLMs used in scientific or regulated decision-making roles.