Skip to main content
llm.info

FineWeb

Open Source
Datasets

15T-token cleaned web dataset

About

FineWeb is Hugging Face's open 15-trillion-token dataset of cleaned, deduplicated Common Crawl web text, with a curated FineWeb-Edu subset. It is a high-quality, permissively licensed corpus for training modern LLMs.

Details

Category
Datasets
License
ODC-By 1.0