Linear probes can identify target variables in Pythia residual streams (160M-12B) as early as step 1,000, yet steering along these directions fails to produce causal effects in most cases. This separation between internal readability and causal efficacy persists across scales.
HOW THIS AFFECTS YOU
●
researcherYou cannot assume that readable internal representations are viable targets for direct mechanistic steering.