Evaluating Physical Experiment Selection Capabilities in Vision Language Models
September 11, 2026
This study evaluates whether VLMs can determine when to stop or request additional physical measurements to solve reasoning tasks. It introduces a benchmark requiring models to choose between immediate answers or selecting the cheapest additional experiment to resolve uncertainty among four possible physical worlds.
HOW THIS AFFECTS YOU
●
researcherYou can use this benchmark to measure active perception and experimental design capabilities in multimodal models.