LAION-BVD provides 80 million video clips totaling 10 million hours of content for multimodal pre-training. The dataset includes synthetically generated video and audio captions derived from content-aware scene detection.
HOW THIS AFFECTS YOU
●
builderYou can leverage this massive, open dataset to pre-train video-text and audio-text multimodal models.
●
researcherThis offers a large-scale resource for studying scaling laws in video-based multimodal learning.