Cross-Lingual Safety Alignment Degrades in Low-Resource Languages
September 24, 2026
Safety alignment transfer across languages fails in low-resource settings when using hard negatives like XSTest contrast prompts. While alignment appears stable with easy negatives, harmfulness representation quality collapses in low-resource languages on models like Qwen2.5-7B-Instruct.
HOW THIS AFFECTS YOU
●
researcherYou should use hard negatives rather than easy negatives to accurately evaluate cross-lingual safety transfer.
●
policyThis indicates a significant safety gap for non-English speaking populations in current LLM deployments.