Study exposes identity-aware bias in LLM-as-a-judge evaluation
August 11, 2026
Research into decentralized LLM evaluation reveals that LLM-as-a-judge methods suffer from identity-aware bias, where models score answers based on their source rather than quality. The study examines this across seven verifier models, including Llama 3.3 70B and DeepSeek V, to address transparency issues in proprietary model benchmarks.
HOW THIS AFFECTS YOU
●
researcherYou must account for source-model bias when using LLMs to evaluate other models.
●
policyThis highlights the risks of relying on unverified, vendor-led benchmarks for model safety and capability claims.