Human Audit Reveals 5-8% Success Underestimation in Web Agent Eval
October 2, 2026
Human review of WebArena-Lite tasks shows that LLM-based evaluators miss 5.45 to 8.49 percentage points of successful task completions. The study also demonstrates that providing procedural guide text and memory support mechanisms improves corrected success rates.
HOW THIS AFFECTS YOU
●
builderImplementing guide text and explicit execution state can improve the reliability of your web agents.
●
researcherYou should account for evaluator bias and false negatives when interpreting web agent benchmark results.