Perfect Aliasing in Truth Probes for Compliant LLMs
September 11, 2026
Research reveals that truth probes in reward-trained models like Gemma-2-9B suffer from perfect aliasing, where they fail to distinguish between truth and prescribed actions. Using mixed-fit probes can solve this, increasing AUROC from 0.006 to 1.000.
HOW THIS AFFECTS YOU
●
researcherThis identifies a critical failure mode in how we probe model internal states for truthfulness.
●
policyThis suggests current methods for detecting deceptive alignment may be fundamentally flawed.