They say:
> in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.”
Does that stack up?
The Meta Paper - https://arxiv.org/pdf/2302.13971.pdf - says:
> We include two book corpora in our training dataset: the Gutenberg Project, which contains books that are in the public domain, and the Books3 section of ThePile (Gao et al., 2020)
The Pile Paper - https://arxiv.org/abs/2101.00027 - says it was trained (in part) on "Books3" which it describes as:
> Books3 is a dataset of books derived from a copy of the contents of the Bibliotik private tracker made available by Shawn Presser (Presser, 2020).
Shawn Presser's link is at https://twitter.com/theshawwn/status/1320282149329784833 and he describes Book3 as
> Presenting "books3", aka "all of bibliotik" - 196,640 books - in plain .txt
I don't have the time and space to download the 37GB file. But if Silverman's book is in there... isn't this a slam dunk case?
Meta's LLaMA is - as they seem to admit - trained on pirated books.