The SF20K dataset provides 3,582 hours of video across 20,143 amateur films, averaging 12 minutes per movie. This enables training and evaluation for story-level video understanding beyond short, single-scene instructional clips.
HOW THIS AFFECTS YOU
●
researcherYou can train vision-language models on much longer temporal contexts and complex narrative structures.