Mid-Network MLP Layers Dominate Safety Refusal Behavior in LLMs
August 13, 2026
Safety-aligned refusal behavior is primarily concentrated in MLP weights rather than attention parameters, with MLP weight transplantation yielding 2.7 times more refusal recovery. Experiments show these refusal-relevant parameters are consistently located in mid-network blocks.
HOW THIS AFFECTS YOU
●
researcherYou can target specific MLP blocks for more efficient safety fine-tuning or weight editing.
●
policyThis finding suggests that safety alignment may be more surgically manipulatable than previously thought.