The Stack v2
Open Source
Datasets
Large permissive source-code dataset
About
The Stack v2 is the BigCode project's open dataset of permissively licensed source code (built with Software Heritage), roughly 4x larger than v1. It is the training corpus behind StarCoder2 and other open code models.
Details
- Category
- Datasets
- License
- Other (per-file licenses)