Small VLMs Require Fine-Tuning to Match Human Cross-Modal Choice Alignment
September 30, 2026
While larger vision-language models show some alignment with human cross-modal associations, their visual saliency often fails to match human gaze patterns. Fine-tuning smaller VLMs on human choice data can bring their decision alignment to human majority-vote levels, though visual attention remains distinct.
HOW THIS AFFECTS YOU
●
researcherYou should distinguish between choice alignment and visual attention when evaluating VLM human-likeness.
●
designerUnderstanding where models look versus humans can help in refining multimodal generative interfaces.