LLM-Based Novelty Evaluation Shows High Sensitivity to Prompting
October 2, 2026
Automated novelty judges suffer from significant instability in scoring generated ideas. Controlled studies using OpenReview data show that minor prompt variations can flip the verdict on more than 50% of identical ideas, complicating the use of LLMs for ideation benchmarking.
HOW THIS AFFECTS YOU
●
researcherYou should exercise caution when using LLMs as zero-shot judges for evaluating creative or novel outputs.