LLM Leaderboard Scores Are Highly Sensitive to Evaluation Harness Configurations
August 25, 2026
Testing 12 open-weight LLMs across 26 harness configurations reveals that benchmark performance is significantly impacted by option order, prompt wording, and decoding methods. The study introduces a fragility grid to identify specific items where configuration variance dictates model ranking.
HOW THIS AFFECTS YOU
●
researcherYou must account for harness configuration when comparing model performance to avoid misleading conclusions.
●
policyThis highlights the need for standardized evaluation protocols to prevent manufactured rankings in regulatory assessments.