Common Crawl Data Stored on a Hugging Face Bucket (commoncrawl.org)

🤖 AI Summary
Common Crawl has announced that its crawl archive, available since April 2026, is now stored on a Hugging Face Storage Bucket, complementing its existing availability on AWS S3. This new integration facilitates easier access to large-scale web data as it leverages the comprehensive tools within the Hugging Face ecosystem, allowing users to browse, download, and interact with data seamlessly. The addition of CDN pre-warming in select regions also enhances read throughput for computations, making it a valuable resource for researchers and developers in the AI/ML domain. The Hugging Face Storage Bucket supports an S3-compatible API, providing flexibility in accessing the data without requiring significant changes to existing pipelines. Various methods, including the Hugging Face CLI and the ability to mount the bucket as a local filesystem, enable users to work with the data effectively. Key features include querying the Common Crawl URL Index using DuckDB without downloading large datasets and utilizing Python scripts to analyze web page data efficiently. These enhancements not only streamline workflows for data scientists and AI practitioners but also extend the usability of expansive web datasets for machine learning applications.
Loading comments...
loading comments...