FineWeb
Open Source
Datasets
15T-token cleaned web dataset
About
FineWeb is Hugging Face's open 15-trillion-token dataset of cleaned, deduplicated Common Crawl web text, with a curated FineWeb-Edu subset. It is a high-quality, permissively licensed corpus for training modern LLMs.
Details
- Category
- Datasets
- License
- ODC-By 1.0