ParaGUIBench is a new benchmark designed to evaluate the parallel execution and coordination of multiple GUI agents on separate desktop instances. It uses a multi-device Docker infrastructure and a 233-task dataset to measure efficiency and step reduction in long-horizon tasks.
HOW THIS AFFECTS YOU
●
builderYou can use this benchmark to evaluate how well your agents scale via multi-agent parallelization.
●
researcherThis provides a necessary metric for moving GUI agents from sequential to coordinated parallel workloads.