Skip to main content
llm.info

C4

Open Source
Datasets

Colossal Clean Crawled Corpus

About

C4 (the Colossal Clean Crawled Corpus) is Google's open, cleaned web-text dataset created for training the T5 models, widely used as a filtered Common Crawl derivative for LLM pretraining and research.

Details

Category
Datasets
License
ODC-By 1.0