ATBench Benchmark Evaluates Multi-Step Agent Safety and Trajectory Risks
August 21, 2026
ATBench introduces a trajectory-level safety benchmark consisting of 1,000 interactions averaging 9 turns and 3.95k tokens. It categorizes agentic risk through a taxonomy of risk sources, failure modes, and real-world harms to capture delayed-trigger risks.
HOW THIS AFFECTS YOU
●
researcherYou can now evaluate agent safety across long-horizon interactions rather than isolated prompts.
●
policyThis framework provides more realistic metrics for assessing the systemic risks of deployed autonomous agents.