Measuring Construct Validity in LLM-as-a-Judge Evaluation
August 26, 2026
This research formalizes LLM-as-a-judge reliability using two independent dimensions: invariance to construct-preserving edits and sensitivity to construct-changing edits. Testing across 7 judges shows that high invariance does not guarantee high sensitivity, meaning current judges may fail to detect subtle but meaningful changes.
HOW THIS AFFECTS YOU
●
builderThis warns against relying solely on simple agreement metrics when using LLMs to grade model outputs.
●
researcherYou should evaluate your automated judges using both invariance and sensitivity metrics to ensure true construct validity.