[arXiv]score: 0.24
Evaluating Perspectival Biases in Cross-Modal Retrieval
September 1, 2026
The 3XCM benchmark quantifies perspectival biases in cross-modal retrieval, where linguistic prevalence and cultural associations override semantic alignment. Findings show image-to-text models favor high-resource languages over semantic accuracy, while text-to-image models exhibit a "tugging effect" where low-resource queries drift toward culturally familiar visual patterns in the joint embedding space.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy