Skip to main content
llm.info

The Stack v2

Open Source
Datasets

Large permissive source-code dataset

About

The Stack v2 is the BigCode project's open dataset of permissively licensed source code (built with Software Heritage), roughly 4x larger than v1. It is the training corpus behind StarCoder2 and other open code models.

Details

Category
Datasets
License
Other (per-file licenses)