Fine-Tuning Creates Blind Spots in Activation Oracles
July 24, 2026
Activation Oracles (AOs) used to probe internal model states are not neutral readouts; they are learned systems shaped by their own training data. In controlled settings, fine-tuning subject models to hide concepts can cause AOs to fail at reading those specific internal representations.
HOW THIS AFFECTS YOU
●
researcherYou must account for the observer effect when using AOs to interpret internal model states.