Modifying 0.19% of LLaMA-2 Weights Compromises On-Device Safety
October 8, 2026
Safety-critical behaviors in LLaMA-2-7B-Chat are concentrated in sparse parameter subsets, specifically the MLP down_proj layer. Modifying just 0.19% of weights in this layer increases Basic ASR to 53% while maintaining 51.6% accuracy on tinyBenchmarks, highlighting significant vulnerabilities for on-device SLMs.
HOW THIS AFFECTS YOU
●
researcherYou can use these localization methods to perform sparse fault analysis on small language models.
●
policyThis demonstrates how easily safety alignment can be bypassed on local devices through targeted parameter manipulation.