The Curved Inference framework extends sleeper agent detection by moving beyond linear probes to analyze semantic complexity and curvature. This methodology simulates realistic deceptive reasoning in multi-turn context windows rather than relying on artificial, easily detectable triggers.
HOW THIS AFFECTS YOU
●
researcherTraditional linear probes may fail to detect sophisticated deceptive alignment that emerges naturally through training.
●
policySafety evaluations must evolve to detect more nuanced, non-linear deceptive behaviors in LLMs.