●builderYou can use this benchmark to validate whether your VLM-based automated evaluators are actually providing reliable feedback for agent training.
●researcherUse this to study the systematic reliability gaps in VLM-as-a-judge paradigms for complex digital workflows.