Audit Reveals Severe Unreliability in LLM-as-a-Judge Frameworks
October 9, 2026
An evaluation of six frontier models shows that LLM judges fail to maintain consistency across identical replications at zero temperature and are highly sensitive to position-order swaps. The study introduces the trustworthy verdict rate (T) to quantify reproducibility, order-invariance, and accuracy.
HOW THIS AFFECTS YOU
●
builderDo not rely on single-pass LLM evaluations as ground truth for your model performance metrics.
●
researcherYou should account for high variance and order bias when using LLMs for automated benchmarking.