C4
Open Source
Datasets
Colossal Clean Crawled Corpus
About
C4 (the Colossal Clean Crawled Corpus) is Google's open, cleaned web-text dataset created for training the T5 models, widely used as a filtered Common Crawl derivative for LLM pretraining and research.
Details
- Category
- Datasets
- License
- ODC-By 1.0