NormViz Benchmark Evaluates Multimodal Cultural Reasoning
September 9, 2026
NormViz-Bench consists of 3,268 contrastive image pairs across 16 countries to test whether VLMs understand local social norms through visual behavior. The benchmark prevents reliance on superficial visual shortcuts by requiring correct classification of culturally relevant actions.
HOW THIS AFFECTS YOU
●
researcherYou can now rigorously evaluate if your multimodal models are biased toward specific cultural norms.
●
policyThis benchmark provides a tool to audit AI systems for cultural sensitivity and global deployment risks.