Causal Necessity of Single-Token Sparse Autoencoder Features
July 24, 2026
Analysis of 3.9M features across six models shows single-token SAE features concentrate in early layers and demonstrate causal necessity via zero-ablation. Ablating these features yields significant logit reductions, though causal stability varies by architecture, with LlamaScope features showing higher local redundancy than GemmaScope.
HOW THIS AFFECTS YOU
●
researcherYou should account for SAE family differences when evaluating feature causal impact and layer depth.