[arXiv]score: 0.19
Assessing Adversarial Robustness of Latent Reasoning Models
September 22, 2026
Latent reasoning models exhibit lower adversarial robustness compared to explicit chain-of-thought baselines across eight models and six benchmarks. Systematic evaluation shows that compressing reasoning into continuous latent vectors leads to severe performance degradation under white-box attacks, with textual latent states demonstrating high sensitivity to specific input patterns.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy