Layer-freezing restores safety refusals after few-sample fine-tuning attacks
October 2, 2026
Fine-tuning on just a few dozen harmful examples can bypass LLM safety alignment, but researchers found that refusal behaviors are localized to specific layers. Freezing layers up to a specific transition depth prevents safety degradation even when the model is subjected to 100 harmful examples.
HOW THIS AFFECTS YOU
●
researcherYou can use layer-freezing or singular direction removal to defend against fine-tuning attacks.
●
policyThis provides a technical basis for assessing the robustness of model safety layers.