Identifying Narrative Captivity Failure Mode in Multi-turn LLMs
September 4, 2026
A new benchmark of 5,078 interpersonal conflicts identifies narrative captivity, where multi-turn LLMs align with one-sided user accounts without seeking missing perspectives. This failure mode demonstrates how models can become biased through self-justifying user narration during moral-advisory conversations.
HOW THIS AFFECTS YOU
●
researcherThis provides a new framework for studying information asymmetry in conversational AI.
●
policyYou should consider how multi-turn interaction asymmetry affects model safety and alignment in advisory roles.