MVVBench Evaluates 4D Reasoning in Vision-Language Models
September 28, 2026
MVVBench is a new benchmark for multi-view video understanding that requires models to integrate spatial and temporal evidence across non-overlapping camera streams. Questions are designed to be unanswerable from any single view or timestamp, forcing joint 4D reasoning.
HOW THIS AFFECTS YOU
●
researcherUse this benchmark to test whether your VLMs can actually reason across time and multiple perspectives.