OTAP evaluates LLM agent performance by treating trajectories as execution graphs rather than binary success flags. It uses an unbalanced fused Gromov-Wasserstein transport problem to provide a distance metric that is invariant to dependency-preserving reorderings.
HOW THIS AFFECTS YOU
●
builderYou can use this metric to better debug why agentic workflows fail beyond simple success/failure outcomes.
●
researcherThis provides a more mathematically rigorous way to benchmark agentic reasoning and planning.