Measuring Resolution Limits of LLM-as-a-Judge Benchmarks
September 24, 2026
Analysis of 373,019 judgments shows that benchmark resolution is capped by the judge's ability to distinguish between systems, regardless of item count. Using pairwise preference instead of pointwise rubric scoring reduces error variance by two orders of magnitude and raises the accuracy ceiling to 0.986.
HOW THIS AFFECTS YOU
●
researcherYou should prioritize pairwise preference testing over pointwise scoring to increase the measurable resolution of your benchmarks.