AgentHorizon Benchmark for Long-Horizon Computer-Use Evaluation
October 9, 2026
AgentHorizon introduces a benchmark of 1,373 tasks based on 166 hours of human-recorded computer usage across three operating systems. It uses instruction-swapping to create negative tasks, specifically testing whether automatic judges can detect subtle constraint violations in long trajectories.
HOW THIS AFFECTS YOU
●
builderYou can use this benchmark to stress-test the automated evaluation pipelines for your computer-use agents.
●
researcherThis provides a more rigorous way to evaluate judge reliability on complex, multi-step trajectories.