An evaluation of 32,874 calls across nine direct-action and six local VLM models reveals widespread failure in visual grounding. Many models remain image-invariant or fail to maintain spatial consistency under lane-axis reflection, with simple constant-speed policies often outperforming complex VLM controllers.
HOW THIS AFFECTS YOU
●
builderDo not trust zero-shot VLM controllers for robotics without rigorous perception-integrity checks.
●
researcherYou should verify visual grounding with ablation tests rather than relying on success trajectories.