SAE Latents Capture Morpho-Syntactic Structure via Feature Groups
September 23, 2026
Analysis of Sparse AutoEncoders (SAEs) reveals that part-of-speech categories are encoded by compact groups of sparse latents rather than one-to-one mappings. These linguistic structures remain stable on held-out data and are not solely driven by lexical memorization.
HOW THIS AFFECTS YOU
●
researcherYou can use these findings to better interpret how linguistic features are distributed across SAE feature spaces.