DeflectBench Evaluates Rhetorical Fallacy Generation in Frontier Models
August 28, 2026
DeflectBench tests 23,990 generations across four frontier models to measure susceptibility to generating whataboutism, ad hominem, and red herring fallacies. Results show model refusal is driven by prompt framing rather than claim content, with certain educational prompts reducing refusal rates to near zero.
HOW THIS AFFECTS YOU
●
researcherYou can use this benchmark to test the robustness of safety training against deceptive reasoning.
●
policyThis highlights how easily safety guardrails can be bypassed through strategic prompt engineering.