Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction
August 12, 2026
Evaluating defensive LLMs via trust-chain localization reveals that intervention rates vary from 0% to 96.3% across five models. The study uses a 300-case housing corpus to test whether defenders identify structural risks like authority or asset control versus surface cues in turn-by-turn interactions.