NEEDLE Removes LLM Backdoors via Weight Orthogonalisation
September 28, 2026
NEEDLE is a training-free method that identifies backdoor directions and refusal subspaces using activation vectors. It applies sequential weight orthogonalisation to suppress triggers without requiring clean reference models or original training data.
HOW THIS AFFECTS YOU
●
builderYou can mitigate backdoor vulnerabilities in existing models without the cost of retraining.
●
policyThis offers a path toward more robust safety alignment through mechanistic interventions.