Multimodal LLM peer review audit reveals high scoring bias
September 1, 2026
Evaluating Qwen2.5-VL-72B and Pixtral-Large-124B on ICLR 2026 submissions shows models score significantly higher than humans, with mean scores between 7.0 and 8.1 versus human averages of 3.4 to 6.8. The models detected only 12.1% of intentionally inserted errors.
HOW THIS AFFECTS YOU
●
researcherNote the significant calibration gap between LLM scores and human expert evaluations.
●
policyThis highlights critical risks in using LLMs for automated academic or professional peer review.