OrchestraBench evaluates multi-agent systems by measuring failure recovery and decomposition quality rather than just task accuracy. Results show intent-reasoning routers significantly outperform keyword-based routers in adversarial enterprise workflows.
HOW THIS AFFECTS YOU
●
builderYou can use these metrics to diagnose why your multi-agent pipelines fail or cascade.
●
researcherYou can move beyond task-accuracy metrics to study agentic recovery and routing robustness.