SAE Features in Gemma 2 and 3 Lack Multilingual Causal Validity
September 7, 2026
Causal validation of Sparse Autoencoder (SAE) features in Gemma 2 and 3 shows that most features recurring across different language settings have small or inconsistent effects on translation behavior. While over 20 features activate frequently across multilingual contexts, they fail to reliably steer or explain cross-lingual model mechanics.
HOW THIS AFFECTS YOU
●
researcherYou should be cautious when assuming that multilingual SAE features represent consistent semantic concepts across languages.