Social Pressure Increases False-Alarm Rates in LLM Safety Panels
August 6, 2026
Simulated peer pressure in multi-model voting panels increases average false-alarm rates from 56.5% to 87.5% when incorrect labels are asserted. The effect is asymmetric, with models following pushes toward unsafe labels at a 75% rate compared to 17% for safe labels.
HOW THIS AFFECTS YOU
●
builderYou can avoid majority voting errors by ensuring peer models do not see misleading consensus prompts.
●
policyYou should be wary of using multi-model voting for safety moderation if the models share the same input context.