Human-LLM Agreement Metrics Fail to Capture Expert-Level Qualitative Coding Quality
August 3, 2026
Blind expert verification reveals that standard human-LLM agreement metrics mischaracterize model performance. In a study of 2,560 messages, human-LLM agreement (Jaccard 0.30) was lower than human-human agreement (0.52), yet expert judges found human and LLM outputs indistinguishable in quality.
HOW THIS AFFECTS YOU
●
researcherYou should reconsider using simple agreement metrics like Jaccard as the sole proxy for LLM qualitative coding accuracy.