Detecting Hidden LLM Behaviors Using Activation-Matched Finetuning
September 2, 2026
Activation-matched finetuning identifies hidden behaviors like backdoors or sandbagging by comparing a suspect model's activations against an anchor model finetuned on benign data. The method detects triggers by measuring high residuals in activation patterns when the suspect model encounters semantic neighbors of hidden behaviors.
HOW THIS AFFECTS YOU
●
researcherYou can use this unsupervised method to detect triggers without prior knowledge of the specific behavior.
●
policyThis provides a new technical framework for auditing models for undisclosed censorship or safety triggers.