Reading an LLM judge's verdict directly from its first token's logits distorts position bias results, acting as an upper bound rather than a true measure. In tests with Qwen3, judges failed to lead with a verdict in up to 49% of cases, causing forced reads to return the first response regardless of actual judgment.
HOW THIS AFFECTS YOU
●
builderYou must ensure your evaluation pipelines allow for full generation to obtain valid model comparisons.
●
researcherYou should avoid using logit-based readout for evaluation to prevent artificially inflated position bias metrics.