MMSAFE Identifies Safety-Degrading Data Across Multilingual Layers
September 22, 2026
MMSAFE is a multi-layer framework designed to detect safety-degrading samples in multilingual fine-tuning data. The method accounts for the fact that safety signals are distributed across different layers depending on the language, rather than being concentrated in a single layer.
HOW THIS AFFECTS YOU
●
researcherThis provides a more robust way to filter training data for multilingual safety alignment.
●
policyYou can better govern the safety of multilingual models by identifying hidden risks in fine-tuning sets.