Introduces a benchmark of 3,862 questions across 182 scenes to evaluate predictive spatial intelligence in VLMs. It moves beyond static perception to test an agent's ability to anticipate scene changes after interventions.
HOW THIS AFFECTS YOU
●
builderThis provides a metric for evaluating agents intended for physical or robotic environments.
●
researcherYou can use this to diagnose why VLMs fail at long-horizon spatial reasoning.