HuRo Dataset Scales VLA Pretraining with Robotized Human Videos
September 17, 2026
The HuRo pipeline converts heterogeneous human videos into robot-aligned observations and action trajectories to bridge the embodiment gap. The resulting dataset contains 630K robotized episodes and 142M frames for scalable Vision-Language-Action (VLA) pretraining.
HOW THIS AFFECTS YOU
●
builderYou can leverage vast amounts of human video data to pretrain robotic control models.
●
researcherThis offers a scalable way to address the data scarcity problem in robot embodiment.