Audit Finds 10.8x Increase in Safety Refusal Failure for Intimate Framing
September 25, 2026
An audit of six conversational AI systems reveals that shifting from non-intimate to intimate-partner descriptors increases non-refusal rates by up to 10.8-fold. While models like ChatGPT 5.2 and Claude Sonnet 4.5 exhibit high refusal rates generally, harmful content frequently leaks through when framed within relational contexts.
HOW THIS AFFECTS YOU
●
researcherThis highlights a critical gap in current refusal logic regarding how prompt framing influences safety thresholds.
●
policyYou should account for relational framing vulnerabilities when setting safety guardrails for conversational agents.