Model Merging Causes Asymmetric Collapse of Safety Capabilities
July 31, 2026
Merging Gemma-3-1B-IT fine-tunes using Linear, SLERP, TIES, or DARE-TIES results in asymmetric capability retention. Merged models maintain 81-85% jailbreak refusal rates while classification accuracy for harm-level detection collapses to at most 12.9%.
HOW THIS AFFECTS YOU
●
builderBe cautious when merging safety fine-tunes, as refusal behaviors may overwrite specific classification capabilities.
●
policyThis demonstrates that merged models may fail to categorize harms correctly even if they refuse to engage with them.