Adversarial Training Decouples Representational Simplicity from Causal Circuit Size
September 30, 2026
Using GPT-2 Small, this study demonstrates that adversarial training can reshape internal representations to be simpler without reducing the actual size of the causal circuits required to perform tasks. This suggests that interpretability metrics like feature attribution do not always predict circuit tractability.
HOW THIS AFFECTS YOU
●
researcherThis warns you that sparse autoencoder simplicity may not directly equate to easier mechanistic interpretability.