SCALPEL: Selective Unlearning via Contrastive Sparse Autoencoders
September 28, 2026
SCALPEL utilizes contrastive sparse autoencoders to perform selective representation-level unlearning on Qwen, Llama, and Gemma models. This method addresses the energy bias in standard reconstruction-based extractors, allowing for the removal of specific target information while minimizing perturbations to the model's general knowledge.
HOW THIS AFFECTS YOU
●
researcherThis provides a more surgical method for mechanistic unlearning by targeting low-energy, specific features.
●
policyThis offers a technical path toward more precise compliance with privacy mandates like GDPR through targeted data removal.