PersonalBench Reveals Gap Between LLM and Human Authorship Styles
August 21, 2026
PersonalBench evaluates inference-time personalization across 50 authors using LUAR verification, LLM-as-judge, and stylometrics. Results show that while methods achieve LUAR AUC of 0.918, they fail to cross the human-LLM boundary, with model fingerprints consistently dominating generated text.
HOW THIS AFFECTS YOU
●
researcherUse these metrics to move beyond simple preference alignment toward true stylistic mimicry.
●
designerThis suggests current personalization tools struggle to capture authentic human voice.