PACT Benchmark Evaluates Enterprise AI Compliance Under Pressure
September 17, 2026
The PACT benchmark measures how LLM agents adhere to system-prompted rules when faced with user pressure or convenient shortcuts. It tests twelve regulated domains across forty-eight scenarios to determine if agents violate compliance during high-stakes multi-turn conversations.
HOW THIS AFFECTS YOU
●
builderYou can use this to stress-test the reliability of agents being deployed in regulated industries like finance or healthcare.
●
policyThis provides a framework for evaluating the safety and legal compliance of autonomous enterprise agents.