LLM Judges Exhibit Anchor Bias in Entity Alignment
October 8, 2026
A systematic study reveals that LLM-as-judge evaluations for entity alignment suffer from anchor bias, where visible system decision labels cause judge discrimination to collapse (J-ROC-AUC 0.12-0.87). Implementing a label-free protocol recovers near-ceiling performance for evaluating structured prediction tasks.
HOW THIS AFFECTS YOU
●
builderAvoid passing system labels to your judge LLM to prevent significant performance degradation in your evaluation pipeline.
●
researcherYou should account for label exposure when designing evaluation frameworks for structured prediction.