Skip to main content
llm.info

Common Crawl

Open Source
Datasets

Petabyte-scale open web crawl

About

Common Crawl is a free, open corpus of web crawl data spanning petabytes and years of the public web. It underpins the pretraining data of most large language models and is available on AWS and via monthly crawl archives.

Details

Category
Datasets
License
Common Crawl Terms