"So here’s the big picture. There are three sets of datasets: 1. Data exists out there in the world. It has been collected into datasets and posted online. I’ll call this raw data. 2. We take that data, clean it, and process it for language modeling. I’ll call this per-set data. 3. We combine those per-set data into one massive dataset, the Pile. This is heavily processed, including weighing the components.
We created 2 and 3 and put them online. We put 2 online so that people can reweigh and remix the data if they wish, but we expect most people to just download 3 and use it out of the box. Access to 3 will be provided in several forms, including HuggingFace and from our website.
2 and 3 are not copyright violations, even if the data is copyrighted, because they fall under fair use (at least in the US).
The Pile contains code that turns 1 into 2 and code that turns 2 into 3.
When you download Maroon 5 from a website, you are creating a dataset corresponding to 2. That can be copyright violation depending on what you do with it, but our use is not a copyright violation."