Live Circuit Extraction via Single Forward Pass Attention Routing
September 23, 2026
This method extracts mechanistic interpretability circuits by treating attention weights as a routing map during a single forward pass. Testing on GPT-2 and Pythia models shows that ablating these extracted edges significantly degrades performance on induction and IOI tasks compared to random ablation.
HOW THIS AFFECTS YOU
●
researcherYou can now perform cheap circuit sketching without the heavy computational cost of full head-by-head patch sweeps.