The Pile: An 800GB Dataset of Diverse Text for Language Modeling
arxiv.org
arxiv.org
As far as I understand RedPajama has 1.2T (https://github.com/togethercomputer/RedPajama-Data) and has a table in the readme listing the main parts and how many tokens each part has.
— anonymous shaper