Transformer Residual Streams Encode Partner Expertise Early in Processing
September 9, 2026
Analysis using the ExpertCollab corpus shows that a model's inference of a dialogue partner's expertise is most decodable in early layers and becomes near-chance by the network midpoint. Counterfactual patching reveals that expertise information injected after the midpoint propagates significantly more effectively than injections at peak decodability layers.
HOW THIS AFFECTS YOU
●
researcherYou can use these findings to better understand the temporal separation between feature encoding and functional utilization in transformer architectures.