Toolspublished

LAION Releases 10-Million-Hour Video Corpus for Research, Excluding Commercial Use

The release brings video, sound and still frames into one open research resource, while synthetic labels, English-heavy sourcing and noncommercial terms narrow how it can be used.

By 2 min read
LAION Releases 10-Million-Hour Video Corpus for Research, Excluding Commercial Use
LAION Releases 10-Million-Hour Video Corpus for Research, Excluding Commercial Use

Listen to this story

The audio brief

About 1:31
0:001:31
Read transcript
LAION has released an enormous video corpus for AI research: 80 million web videos totaling 10 million hours, bundled with sound and still images. The release, called the Big Video Dataset, or BVD, started with 1.3 billion video URLs gathered from Common Crawl. LAION downloaded and processed 80 million of them, then used scene detection to divide the footage into 55 million clips. Each clip comes with synthetic captions for both its video and audio, giving researchers a single multimodal resource for connecting what appears on screen with what is heard and described in text. There is also a substantial image component. LAION extracted 300 million frames, including frames around scene changes, and says their visual mix differs from standard web-image datasets. That could make the corpus useful for image-text training as well as video and audio work. LAION reports that ViCLIP models trained on BVD matched or beat models trained on InternVid by as much as 2.1 percent on standard video-text benchmarks, as training expanded from 10 million to 50 million clips. Those results are LAION’s evaluations, and the practical limits are significant: BVD is free for research, but commercial use is excluded. Most of the material comes from YouTube, and most is in English, so inherited bias and weak coverage of other regions and languages remain concerns. The key constraint is whether researchers can use this scale while still respecting creator rights, platform terms, and applicable law.

Story brief

3 key points

LAION’s August 29 release turns 80 million downloaded web videos into a multimodal research corpus spanning 10 million hours, 55 million captioned clips, and 300 million extracted frames. The scale could lower data barriers for video, audio, and image-model researchers, while reported benchmark gains over InternVid remain limited to LAION’s evaluations. BVD is freely available but barred from commercial use, and its...

  1. 01

    LAION began with 1.3 billion Common Crawl video URLs and processed 80 million into BVD.

  2. 02

    Content-aware scene detection produced 55 million clips with synthetic video and audio captions.

  3. 03

    LAION reports ViCLIP matched or exceeded InternVid by up to 2.1% as training scaled from 10 million to 50 million clips.

LAION has released a research-only corpus of 80 million web videos totaling 10 million hours. The Big Video Dataset, or BVD, gives researchers one resource for training across video, audio and images, while excluding commercial use.

From web crawl to training corpus

The project began with 1.3 billion platform-specific video URLs collected from Common Crawl. LAION downloaded and processed 80 million of them into LAION-BVD, releasing the collection for AI research on August 29.

The raw footage is only part of the release. LAION used content-aware scene detection to make clips, then synthetically generated video and audio captions. The result includes 55 million captioned clips, designed to connect visual content, sound and text during training.

BVD at a glance
1.3 billionVideo URLs collected

LAION collected 1.3 billion platform-specific video URLs from Common Crawl.

80 millionVideos downloaded

BVD contains 80 million downloaded videos.

10 million hoursCombined duration

Those videos have a combined duration of 10 million hours.

A second use for the footage

LAION also extracted 300 million video frames for image-text pre-training. It says the scene-changing frames have a visual distribution distinct from standard web-image corpora, positioning the dataset as a still-image source as well as one for motion and sound.

LAION says ViCLIP models trained on BVD matched or exceeded InternVid-trained models by up to 2.1% on standard video-text benchmarks as training grew from 10 million to 50 million clips. It also reports competitive audio-text results and strong image-text retrieval performance.

Terms and limits remain part of the release

LAION describes BVD as infrastructure for open and reproducible multimodal research, including safety analysis. The dataset and code are freely available, but the release is exclusively for research and not licensed for commercial use.

The corpus has visible representation limits: most videos come from YouTube, and a majority are in English. LAION warns that BVD may contain biases, stereotypes and uneven representation across languages, regions and topics, which models trained on it may inherit.

LAION asks users to respect creators’ rights, applicable laws and platform terms. The Decoder says LAION can likely point to a 2024 Hamburg Regional Court ruling that allowed it to collect copyrighted content for non-commercial research.

Sources

  1. the-decoder.comLAION drops massive open video dataset with 10 million hours of footage for AI research
  2. projects.laion.aiLaion Big Video Dataset