Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents
September 15, 2026
PAI-Bench evaluates agent identity through a provider-neutral benchmark separating factual recall from behavioral enactment and resistance to instruction drift. Testing sixteen synthetic profiles revealed that explicit field cues increased the joint presence of identity identifiers from 0/8 to 7/8. The framework uses external scoring oracles to audit identity fidelity across versioned updates.