Common Crawl hosts datasets on Hugging Face for easier access
September 16, 2026
Common Crawl data is now accessible via Hugging Face buckets, providing direct access to web graphs and CDXJ indexes. This integration simplifies the data ingestion pipeline for training large-scale web-scale language models.
HOW THIS AFFECTS YOU
●
builderYou can pull large-scale web datasets directly into your training workflows without managing complex scraping infra.
●
researcherThis reduces the friction of acquiring massive web corpora for pre-training and evaluation.