The M2BIND benchmark reveals that vision-language model concept binding is not language-invariant. Cross-family and cross-script language shifts trigger binding collapse, where internal associations lose causal strength and shift to later layers.
HOW THIS AFFECTS YOU
●
builderYou cannot assume global performance stability for VLMs based solely on monolingual benchmarks.
●
researcherThis highlights a critical failure mode in how VLMs handle cross-lingual semantic associations.