OSReward: Benchmarking VLM Reliability for Computer-Use Agents
July 29, 2026
OSReward is a benchmark designed to evaluate the reliability of Vision-Language Models (VLMs) acting as judges for computer-using agent (CUA) trajectories. It uses diverse agent backbones and human-verified instructions across multiple platforms.
HOW THIS AFFECTS YOU
●
builderYou should verify your reward models against this benchmark before relying on them for agent RL.
●
researcherThis highlights the potential unreliability of using VLMs for automated reward modeling.