TANGO Whole-Body Vision-Language-Action Model for Humanoids
September 7, 2026
TANGO is a vision-language-action framework that predicts 29-DoF joint-space actions for humanoid robots navigating cluttered 3D environments. It enables continuous geometry-aware whole-body adaptation using egocentric RGB observations and natural language instructions.
HOW THIS AFFECTS YOU
●
builderYou can implement complex humanoid navigation that coordinates arms, torso, and gait for obstacle avoidance.
●
researcherThis moves beyond 2D path planning to continuous 3D whole-body control in simulation.