Atomic Feature Theory Validated via Scaling Sparse Autoencoders
October 6, 2026
A new theory of language model representations suggests that sparse autoencoders (SAEs) recover an increasing prefix of prevalent atoms as they scale. Testing on SAE sizes from 512 to 131,072 demonstrated that larger SAEs recover both parent and child features, contradicting the belief that features merely split during scaling.
HOW THIS AFFECTS YOU
●
researcherYou can treat scaling SAEs as a viable path toward recovering a more complete and stable dictionary of model features.