LongRCA Bench provides 1,140 failed trajectories to help developers localize the exact step and role responsible for agent failures in long-horizon tasks. The benchmark features trajectories with a median length of 145 steps, targeting the current difficulty in error attribution.
HOW THIS AFFECTS YOU
●
builderYou can use this benchmark to improve the debugging and observability of your long-running agentic workflows.
●
researcherThis provides a more rigorous way to evaluate how models handle error localization in complex sequences.