Discovery of Recognition-Refusal Misalignment in LLMs
August 28, 2026
Research shows LLMs represent the impossibility of certain math and code queries in a hidden state direction that is nearly orthogonal to the direction used for safety-refusal.
HOW THIS AFFECTS YOU
●
researcherYou can investigate the decoupling of logic recognition and safety routing.
●
policyYou should account for the fact that models may 'know' a prompt is invalid but fail to refuse it.