Where did the ambiguity go? Examining how multimodal models interpret polysemous words
August 4, 2026
Title: Where did the ambiguity go? Examining how multimodal models interpret polysemous words
Source: arxiv
Multimodal models exhibit significantly lower semantic variety than humans or text-only models when processing context-free polysemous words. Testing 17 text-to-image and 15 text-generation models shows normalized entropy dropping from 0.47 in human imagination to 0.25 in text and 0.10 in image generation.