Analyzing Consistency and Specificity in Mechanistic Interpretability
September 4, 2026
A study of model circuits reveals that component-level circuits (attention heads/MLP blocks) are highly consistent but lack task-specificity. Conversely, neuron-level circuits show higher task-specificity but suffer from low intra-task consistency.
HOW THIS AFFECTS YOU
●
researcherThis clarifies why component-level ablation often affects multiple tasks simultaneously and limits interpretability gains.