REDAgentBench Automates Red Teaming for LLM Agent Systems
August 12, 2026
REDAgentBench is an executable framework that evaluates agent safety by running attacks in isolated service sandboxes. It measures actual harmful effects through service receipts and state changes rather than relying on simple attack success rates, uncovering a 65.69% macro-average ASR across tested models.
HOW THIS AFFECTS YOU
●
builderYou can test your agent's robustness against tool-use exploits in a sandboxed environment.
●
policyYou can use more accurate measurements of agent vulnerability that account for actual environmental impact.