LAION has launched the Big Video Dataset (BVD), a new open-source collection of video footage aimed at advancing AI research. The dataset includes 10 million hours of video, sourced from over 1.3 billion video URLs collected from CommonCrawl. The team downloaded 80 million of these videos, extracting 55 million clips with auto-generated video and audio descriptions, along with 300 million still images. Most of the footage comes from YouTube, with the majority in English. The dataset is intended for research purposes only and is freely available to the public. Source: thedecoder
According to the paper, models trained on BVD outperform comparable models trained on InternVid by up to 2.1 percentage points on common video-to-text benchmarks. The training process links video, audio, and text, enabling the model to learn which visual content corresponds to specific descriptions or sounds. This approach is designed to improve the accuracy and efficiency of AI systems in understanding and generating video-related content. Source: thedecoder
LAION emphasized that the dataset is for research use only and that users must respect the rights of original content creators. The organization cited a 2024 Hamburg Regional Court ruling that permitted the use of copyrighted content for non-commercial research. The dataset and associated code are freely available to the public. Source: thedecoder