Modality Gap Reduction Can Cause Prediction-Level Hubness in CLIP
September 2, 2026
Reducing the image-text modality gap in CLIP does not guarantee better zero-shot accuracy and can lead to prediction-level hubness. This failure mode occurs when gap correction alters decision margins, causing model predictions to concentrate on a narrow subset of classes.
HOW THIS AFFECTS YOU
●
researcherYou should account for decision margin shifts rather than just average alignment metrics when optimizing multimodal embeddings.