●builderYou should be skeptical of single-scalar success rates when evaluating your agent's performance on existing benchmarks.
●researcherThis highlights a critical need for more robust, non-scripted evaluation oracles in computer-use agent research.