NovGauge Benchmark for Fine-Grained LLM Paper Novelty Assessment
September 11, 2026
NovGauge provides a human-anchored benchmark of 619 paper pairs to diagnose LLM performance in assessing research novelty across task, problem, and method dimensions. Testing reveals hallucination rates between 0% and 39% among evaluated models during novelty diagnosis.
HOW THIS AFFECTS YOU
●
researcherYou can use this diagnostic pipeline to identify specific dimensions where models fail in peer-review tasks.
●
policyYou should remain cautious about using LLMs for automated academic peer review due to high hallucination rates.