MELLA Dataset Improves Cultural Groundedness in Low-Resource Multimodal Models
July 24, 2026
MELLA introduces a dual-source dataset for eight low-resource languages, separating linguistic fluency from cultural visual-textual alignment. Fine-tuning MLLM backbones on this data mitigates cultural hallucinations by providing native web image-alt-text pairs alongside high-quality translated descriptions.
HOW THIS AFFECTS YOU
●
builderYou can use this dataset to reduce cultural hallucinations when deploying MLLMs in non-English speaking markets.
●
researcherThis provides a way to decouple linguistic capability from cultural knowledge in multimodal evaluation.