GUI-Primitives benchmark reveals 32% accuracy for vision-language agents
August 21, 2026
The GUI-Primitives benchmark uses 994 contrastive instruction pairs to isolate spatial reasoning failures in GUI grounding. Testing nineteen vision-language models shows a maximum strict point-in-box accuracy of only 32% across seven spatial relations.
HOW THIS AFFECTS YOU
●
builderYou should expect significant difficulty when relying on VLMs for precise UI automation.
●
researcherThis provides a more granular way to diagnose why agents fail at spatial grounding.