●builderYou can leverage these scenarios to red-team your agents' ability to handle complex, multi-step tool interactions.
●researcherYou can use this arena to test agent safety against cumulative, long-horizon risks rather than static prompts.
●policyThis provides a more rigorous standard for evaluating the safety of autonomous agents with tool access.