LLM Judges Exhibit Accuracy-Dependent Lenience and Bias
September 14, 2026
LLM judges show a strong correlation between their own task accuracy and judging accuracy, yet more capable models receive systematically more lenient scores. The study proposes calibrated weighted majority voting (WMV) to mitigate these biases by weighting judges based on estimated false-positive and false-negative rates.
HOW THIS AFFECTS YOU
●
builderDo not rely on single-model LLM-as-a-judge scoring for high-stakes model evaluation without calibration.
●
researcherYou can use weighted majority voting to calibrate ensemble evaluations against capability-dependent lenience.