Benchmark Reveals Answer Biases in LLM Probabilistic Reasoning
July 31, 2026
A new benchmark of 14,320 procedurally-generated prompts tests 29 models on logical inference over gradable epistemic modals like probably and must. Results show most models exhibit systematic Yes/No biases that are independent of the actual logical form required.
HOW THIS AFFECTS YOU
●
researcherYou should account for surface-level pattern matching when evaluating logical competence.