TAMP-Nav Aligns VLMs with 3D Navigation via 2D Prompting
August 17, 2026
TAMP-Nav introduces a Pixel-to-3D Action Formulation that reformulates embodied navigation into 2D visual prompting. This allows Vision-Language Models to select 2D pixels that are projected into 3D coordinates for a low-level SLAM controller.
HOW THIS AFFECTS YOU
●
builderYou can build more efficient embodied agents by leveraging a VLM's 2D strengths to control 3D environments.
●
researcherThis method addresses the misalignment between VLM pre-training priors and 3D action spaces.