Concept-Targeted Attribution for Interpreting Linear Probes
August 31, 2026
Concept-Targeted Attribution (CTA) trains attribution graphs to explain specific linear probe directions rather than next-token probabilities. This framework demonstrates that probe-specific circuits can predict probe accuracy with a correlation of rho = 0.91.
HOW THIS AFFECTS YOU
●
researcherYou can use CTA to identify the specific internal circuits that enable a probe to represent a particular concept.