●builderUse this benchmark to test the capability of your VLMs to move beyond description into active 3D environment manipulation.
●researcherThis highlights the gap between textual scene description and precise geometric action in current VLM architectures.
●designerThis signals that high-fidelity 3D interactive AI experiences are still limited by current VLM performance.