FriendBench Evaluates Multimodal Social Inference in LLMs
August 3, 2026
FriendBench benchmarks 26 models on their ability to infer familiarity from 20-second dyadic video, audio, and text clips. While top models match human accuracy, they exhibit a significant prior bias toward labeling pairs as strangers, whereas humans maintain balanced predictions.
HOW THIS AFFECTS YOU
●
researcherYou can use these stimuli and human ratings to evaluate social intelligence and multimodal alignment in your models.