Lie Detection Probes Fail Under Anti-Factual Role-Play Personas
October 1, 2026
Evaluating eight existing lie detection probes on a dataset of 8,916 human-reviewed responses shows they often fail when models adopt anti-factual personas. Probes frequently struggle to distinguish between falsehoods and the internal beliefs of a specific persona, often relying on spurious correlations.
HOW THIS AFFECTS YOU
●
researcherYou should account for persona-driven belief shifts when designing internal state probes.
●
policyThis highlights risks in relying on internal probes for automated truthfulness or safety monitoring.