Authorship Gap in LLM Personalization via Theory-Grounded Metrics
August 17, 2026
Current LLM personalization methods fail to match human stylistic identity, a gap often obscured by uncalibrated evaluation metrics. Using LUAR, a metric grounded in authorship verification theory, researchers found all tested inference-time methods scored below the cross-author baseline of 0.626.
HOW THIS AFFECTS YOU
●
builderCurrent stylistic fine-tuning and prompting techniques are significantly underperforming compared to human authorship ceilings.
●
researcherGround your personalization benchmarks in authorship science rather than ad hoc LLM-as-a-judge methods.