Distributed Part-of-Speech Representations in Sparse AutoEncoder Latents
September 25, 2026
Morpho-syntactic information in language models is encoded in compact groups of sparse latents rather than atomic, one-to-one mappings. Analysis shows that part-of-speech categories are highly recoverable from SAE activations and remain stable on held-out data.
HOW THIS AFFECTS YOU
●
researcherYou should expect to find linguistic structures in SAEs as distributed feature groups rather than single latent neurons.