LLMs Fail Negation Tasks by Repeating Positive Answers
October 8, 2026
LLMs fail negation benchmarks 37-71% of the time by repeating the original answer instead of excluding it. Mechanistic analysis shows specialized attention heads and MLP neurons attempt to suppress original answers and promote candidates, a process that differs fundamentally from human cognitive negation patterns.
HOW THIS AFFECTS YOU
●
researcherYou should consider how non-human negation mechanisms might lead to unexpected reasoning failures in your evaluations.