"Books3 is a dataset of books derived from a copy of the contents of the Bibliotik private tracker made available by Shawn Presser (Presser, 2020). Bibliotik consists of a mix of fiction and nonfiction books and is almost an order of magnitude larger than our next largest book dataset (BookCorpus2). We included Bibliotik because books are invaluable for long-range context modeling research and coherent storytelling"
“They’re not books, man, they’re a dataset!”
(/s)
Open source is generous sharing, and ethical. Start to nibble away at those ideals and at what point do you slip into unethical? IMDB was a crowd-sourced database put together by a wide community pitching in small efforts which one guy was maintaining like an FAQ. Then the guy maintaining it said, "It's worth money, I own it, screw all of you." How would people react if this happened to wikipedia? But wikipedia is safe because it's a non-profit... you know, like OpenAI, right?
Fair to say that whether or not this is correct is pretty important to all the outstanding court cases on this matter.
But you still have to get your hands on the copyrighted data legally. It might be legal to scan every book an institution owns, and train off it, so long as those scans are not distributed. But it is probably not legal to scrape copyrighted content off torrents - creating the copy to train with is infringing, even if the model's final product maybe isn't.
[1] https://fairuse.stanford.edu/overview/fair-use/four-factors/
books.google.com has been allowed to copy all the books they can lay their hands on, so long as they don't regurgitate them in full, so it's not really the taking, but any subsequent reproductions. And the effect on the market is insubstantial if the alternative wasn't going to be the equivalent sales.
If you take every paragraph in the Harry Potter saga and sort the paragraphs in alphabetical order, it's just as good for training short-context-window models, but not a "harm to the market" leading to a lost sale for anyone who wants to read the books.
I'm pretty sure that's what the lawyers will argue in the Silverman case for example. It's going to be interesting to see how the courts decide.
In the present moment on the other hand, it is the entities in the AI industry (e.g. MS) that have the money and can hire the lawyers to convince the judges. Realistically speaking, it's very likely that things will swing the way of AI companies, which will benefit, albeit indirectly, these guys, even though by themselves they're too small to push their agenda, they're just bit players.
https://the-eye.eu/public/Books/ThoseBooks/Puzzles.tar -- 20-Jan-2023 14:54 -- 6M
and it pretends to be a jigsaw puzzle, but is actually eISBN 9781594868573 - The South Beach diet cookbook / Arthur Agatston