Localizing LLM Backdoors Using Sparse Autoencoders
September 9, 2026
Researchers used sparse autoencoders (SAEs) to identify trigger-relevant features in 1B and 8B parameter models during language-switching tasks. The study finds that while SAE features can detect triggers with near-perfect F1 scores, they do not always control the resulting behavior.
HOW THIS AFFECTS YOU
●
researcherSAEs offer a path toward mechanistic interpretability for backdoor detection.
●
policyUnderstanding backdoor localization is critical for developing verifiable AI safety standards.