UniNav Unified Diffusion Model for Visual Navigation
August 5, 2026
UniNav employs a single diffusion process to jointly denoise future visual observations and continuous waypoint trajectories. The architecture incorporates geometry-aware camera tokens and trains on both labeled trajectories and unlabeled video data to improve spatial grounding.
HOW THIS AFFECTS YOU
●
builderYou can leverage this method to build agents that combine visual foresight with action generation.
●
researcherThe unified transformer approach for visual and action tokens presents a new way to model embodied agents.