Video2Skill: From Streaming Experience to Reusable Embodied Skills
September 28, 2026
Video2Skill introduces a Streaming Embodied Skill Discovery (SESD) benchmark to evaluate whether Vision-Language Models can extract reusable manipulation skills from sequential video streams. The framework moves beyond single-event description by requiring models to maintain a persistent skill library that informs future decision-making in embodied tasks.