Common Crawl
Open Source
Datasets
Petabyte-scale open web crawl
About
Common Crawl is a free, open corpus of web crawl data spanning petabytes and years of the public web. It underpins the pretraining data of most large language models and is available on AWS and via monthly crawl archives.
Details
- Category
- Datasets
- License
- Common Crawl Terms