LAION Releases 10-Million-Hour Video Corpus for Research, Excluding Commercial Use
The release brings video, sound and still frames into one open research resource, while synthetic labels, English-heavy sourcing and noncommercial terms narrow how it can be used.
Listen to this story
The audio brief
Story brief
3 key pointsLAION’s August 29 release turns 80 million downloaded web videos into a multimodal research corpus spanning 10 million hours, 55 million captioned clips, and 300 million extracted frames. The scale could lower data barriers for video, audio, and image-model researchers, while reported benchmark gains over InternVid remain limited to LAION’s evaluations. BVD is freely available but barred from commercial use, and its...
- 01
LAION began with 1.3 billion Common Crawl video URLs and processed 80 million into BVD.
- 02
Content-aware scene detection produced 55 million clips with synthetic video and audio captions.
- 03
LAION reports ViCLIP matched or exceeded InternVid by up to 2.1% as training scaled from 10 million to 50 million clips.
LAION has released a research-only corpus of 80 million web videos totaling 10 million hours. The Big Video Dataset, or BVD, gives researchers one resource for training across video, audio and images, while excluding commercial use.
From web crawl to training corpus
The project began with 1.3 billion platform-specific video URLs collected from Common Crawl. LAION downloaded and processed 80 million of them into LAION-BVD, releasing the collection for AI research on August 29.
The raw footage is only part of the release. LAION used content-aware scene detection to make clips, then synthetically generated video and audio captions. The result includes 55 million captioned clips, designed to connect visual content, sound and text during training.
LAION collected 1.3 billion platform-specific video URLs from Common Crawl.
BVD contains 80 million downloaded videos.
Those videos have a combined duration of 10 million hours.
A second use for the footage
LAION also extracted 300 million video frames for image-text pre-training. It says the scene-changing frames have a visual distribution distinct from standard web-image corpora, positioning the dataset as a still-image source as well as one for motion and sound.
LAION says ViCLIP models trained on BVD matched or exceeded InternVid-trained models by up to 2.1% on standard video-text benchmarks as training grew from 10 million to 50 million clips. It also reports competitive audio-text results and strong image-text retrieval performance.
Terms and limits remain part of the release
LAION describes BVD as infrastructure for open and reproducible multimodal research, including safety analysis. The dataset and code are freely available, but the release is exclusively for research and not licensed for commercial use.
The corpus has visible representation limits: most videos come from YouTube, and a majority are in English. LAION warns that BVD may contain biases, stereotypes and uneven representation across languages, regions and topics, which models trained on it may inherit.
LAION asks users to respect creators’ rights, applicable laws and platform terms. The Decoder says LAION can likely point to a 2024 Hamburg Regional Court ruling that allowed it to collect copyrighted content for non-commercial research.
Sources
- the-decoder.comLAION drops massive open video dataset with 10 million hours of footage for AI research
- projects.laion.aiLaion Big Video Dataset