Scientific-Judgment Collapse in AI-Generated Peer Reviews
September 16, 2026
Training models on synthetic, AI-generated scientific reviews leads to scientific-judgment collapse, characterized by compressed rating distributions and reduced semantic diversity. The study demonstrates this recursive degradation using Llama 3.1 8B fine-tuned on ICLR data.
HOW THIS AFFECTS YOU
●
researcherYou should be cautious about using synthetic review data in training sets for evaluation models.
●
policyThis highlights the risks of allowing AI-generated content to degrade the quality of scientific gatekeeping.