DashArena evaluates LLM capabilities in generating interactive analytic dashboards by requiring systems to produce both a UI and a replayable interaction trajectory. The framework uses a browser executor and a VLM judge, distilled into a DashJudge-8B model, to rank performance using Bradley-Terry aggregation.
HOW THIS AFFECTS YOU
●
builderYou can use DashJudge-8B to evaluate how well your agents handle complex data exploration tasks.
●
researcherYou can use a task-grounded benchmark that moves beyond static visual evaluation.