Evaluation of seven LLMs shows they fail to align with human social pragmatics, frequently overproducing neutral labels and underpredicting impolite ones. This systematic bias suggests models struggle with rapport-building strategies and complex linguistic cues.
HOW THIS AFFECTS YOU
●
researcherYou should account for neutral compression when using LLMs as evaluators for social alignment.
●
designerModel-generated social interactions may feel unnaturally neutral or fail to capture nuanced human rapport.