A study of 9,312 judgments across Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5 reveals a same-family lift of 3.4-8.4 percentage points in pairwise evaluations. The researchers propose a corrected estimator to decouple judge family preference from actual candidate quality.
HOW THIS AFFECTS YOU
●
builderYou should account for family-conditioned preference when using LLMs to evaluate models from the same developer.
●
researcherThis provides a method to correct for systematic biases in LLM-based evaluation benchmarks.