Classical Test Theory Limitations in LLM Judge Evaluation
September 25, 2026
Applying classical test theory statistics like KR-20 and the dependability index to LLM judges produces misleading reliability metrics. The research shows these coefficients fail to separate judge error from item design, making them unreliable for characterizing judge performance.
HOW THIS AFFECTS YOU
●
builderBe cautious when using automated evaluation benchmarks that rely on traditional psychometric reliability scores.
●
researcherYou should avoid using standard reliability coefficients to claim LLM judge consistency.