Agentic Computer Use Performance Caps at 20.6% on OSWorld-V2
July 30, 2026
Current LLM agents struggle with long-horizon computer tasks, achieving only 20.6% completion on OSWorld-V2 and 26.2% on Agents' Last Exam. Increasing success rates incurs sharply rising token costs, making autonomous software interaction currently unreliable and expensive for production workflows.
HOW THIS AFFECTS YOU
●
builderYou should expect high latency and low reliability if building agentic workflows for complex GUI tasks.
●
researcherCurrent benchmarks highlight a significant gap between chat performance and functional computer interaction.