Formatting Sensitivity in LLM Judges Increases False Acceptances
October 9, 2026
Small changes in request presentation, such as adding a single colon, increased false acceptance rates from 1% to 26% on GSM8K and DROP benchmarks. The study demonstrates that structured JSON keys and candidate presentation significantly impact judge reliability.
HOW THIS AFFECTS YOU
●
builderYou should carefully design prompt structures to prevent automated judges from overlooking errors in your agent's output.
●
researcherYou must account for presentation bias when evaluating LLM-as-a-judge performance.