Horizon Residual Metric for Long-Horizon Agent Evaluation
July 31, 2026
Long-horizon agent failures are often caused by trajectory-induced degradation where earlier errors compound. The paper proposes the horizon residual—the log-ratio between actual full-task success and a baseline prediction from short, individual stages—to isolate true long-horizon performance.
HOW THIS AFFECTS YOU
●
builderYou can better diagnose why agents fail as conversation history and tool outputs accumulate.
●
researcherYou can use the horizon residual to differentiate between task complexity and context rot.