VA-Bench Evaluates Embodied Spatial Intelligence via Active Perception
September 16, 2026
VA-Bench assesses the observe-reason-act-revise loop in MLLMs using 14 task families across single and dual-arm setups. It requires models to learn procedural context from RGB demonstrations, select camera viewpoints, and issue metric Cartesian commands without privileged object poses or oracle trajectories.
HOW THIS AFFECTS YOU
●
researcherYou can use this benchmark to evaluate how well your MLLMs handle closed-loop spatial reasoning and error correction.