AMRA Defends Against Abliteration via Refusal Aliases
August 20, 2026
AMRA mitigates abliteration by applying rank-k updates to residual stream writer matrices and replacing refusal-inducing activations with random aliases. On Llama-3-8B, this method improves post-abliteration refusal scores by 2.16 points while maintaining MMLU performance within 0.5 percentage points.
HOW THIS AFFECTS YOU
●
researcherYou can explore how weight-editing techniques can stabilize alignment against projection-based attacks.
●
policyThis indicates that post-training alignment can be hardened against weight-orthogonality removal techniques.