CHIARO benchmark evaluates LLM performance on contrastive emotion inference
September 4, 2026
The CHIARO benchmark introduces 1,000 human-annotated sentences to test how models handle single events that trigger opposing emotions in different people. Frontier LLMs achieved a 67.3 macro-F1, significantly trailing human agreement, while standard classifiers performed near chance.
HOW THIS AFFECTS YOU
●
researcherYou can use this benchmark to evaluate model sensitivity to complex, multi-perspective social contexts.
●
designerThis highlights a current limitation in how AI perceives nuanced emotional dynamics in user interactions.