MOSAIC Benchmark for Multimodal Theory of Mind Action
August 24, 2026
The MOSAIC benchmark evaluates whether vision-language models can translate mental state inferences into coordinated verbal and nonverbal behaviors. Results from 13 models show that current VLMs fail to produce spatial or expressive behaviors consistent with expected social outcomes under Theory of Mind constraints.
HOW THIS AFFECTS YOU
●
researcherThis exposes a significant gap between social reasoning capabilities and actual behavioral execution in VLMs.
●
designerExpect current multimodal agents to struggle with nuanced, non-verbal social coordination in interactive environments.