LLM-as-a-Judge Fails to Align with Human Creativity Assessments
August 26, 2026
Experiments show that LLM-based judges exhibit systematic bias toward AI-generated text, favoring stylistic consistency over the unpredictability found in human-authored stories. Automatic metrics and LLM evaluations demonstrate significant misalignment with human judgments across 11 dimensions of creativity.
HOW THIS AFFECTS YOU
●
researcherAvoid relying on LLM-as-a-judge for evaluating creative or non-standard text generation.
●
designerCurrent automated tools may favor generic AI aesthetics over true creative variety.