Syntactic Sensitivity in LLMs Undermines Safety Alignment
August 7, 2026
Research shows that non-imperative syntactic forms can trigger jailbreaks in 16 models up to 70B parameters. Causal mediation analysis reveals that refusal behaviors are conditioned on upstream syntactic features, often due to linguistically biased post-training data.
HOW THIS AFFECTS YOU
●
builderYou need to increase syntactic diversity in training data to prevent grammar-based safety bypasses.
●
policyThis identifies a structural vulnerability in how alignment is currently achieved through post-training.