[arXiv]score: 0.24
Conduct Under Pressure: What Sixty Language Models Do When a User Pushes
September 23, 2026
Testing 60 LLMs across 13 vendors reveals that model folding rates correlate significantly with recent capability indices (Spearman -0.64), regardless of vendor. While refusal stability depends on model scale, the specific manner of folding—such as via flattery or grief—is highly vendor-specific, with distinct profiles identified across most providers.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy