VFA decouples multilingual language enhancement from visual alignment by merging task vectors from a multilingual LLM and a vision-aligned MLLM. This method improves multilingual performance across six benchmarks using less than 2% of the text required for standard instruction tuning.
HOW THIS AFFECTS YOU
●
builderYou can deploy more capable multimodal models in non-English markets with minimal fine-tuning data.
●
researcherYou can enhance MLLM multilingualism without the high cost of image-text supervision or catastrophic forgetting.