MultihopSpatial Benchmark for Multi-Hop VLM Spatial Reasoning
September 7, 2026
MultihopSpatial introduces a benchmark for 1- to 3-hop compositional spatial reasoning in Vision-Language Models. It utilizes a new Acc@50IoU metric to simultaneously evaluate reasoning accuracy and precise visual bounding box grounding.
HOW THIS AFFECTS YOU
●
builderThis provides a more rigorous way to test the spatial grounding of VLA agents intended for physical environments.
●
researcherYou can use this to evaluate how well VLMs handle complex, multi-step spatial queries.