Energy Landscape Analysis Explains Diffusion Language Model Jailbreaks
September 28, 2026
Jailbreaks in diffusion-based LLMs succeed by either obscuring safety disposition at initialization or intervening mid-trajectory to bypass denoising energy barriers. The authors propose three training-free detection signals based on initial logit distributions and trajectory velocity to identify these attacks.
HOW THIS AFFECTS YOU
●
researcherUse these kinetic energy signals to develop more robust detection mechanisms for diffusion models.
●
policyThis provides a mathematical framework for understanding how safety alignment can be bypassed.