AI Books4 Dataset for training LLMs further
old.reddit.com
old.reddit.com
Should they pay too?
If this was enforced would there be search engines?
I am sure there are wild times ahead. At some time it will become outlawed in the way regular piracy is. Not because of book authors or publishers, but either because of Hollywood or AI safety being used by bigtech to limit competition.
What I can grasp is "output of copyrighted data is violation of copyright", and the fix seems to be a straight forward dumb software contentID filter on the output.
One of the major dividing lines between legal and illegal is whether it's a purely mechanical process, or done by human hands. The training of LLMs and diffusion models is a purely mechanical process, more like photocopying or photographing than an artist gaining inspiration"
That's part of where I think a lot of confusion is - LLM training doesn't do any of that. There is no coherent corpus of data in these models.
The argument is that the 400k books being shared here are themselves copyrighted works that the poster probably does not have the right to be distributing, and that is not very nice of them.
"It's just a dataset, bro!" being used to bulldoze people's rights to their own work might have negative consequences.
What's the rate of forgetting for the ingested material for human brain?
I think it's 0. Maybe a handful if pirate libraries of readable ebooks and papers didn't already exist, but even a handful of copyright violators wouldn't be a serious commercial threat to the publishers.
So far, governments have been unable to do anything about pirate libraries. Complaining about AI datasets, poorly formatted for human consumption, seems misplaced.
Criticizing this as "stealing" or "piracy" is a vote for a future where only big tech, or the major publishers themselves, can train models on good datasets. Nobody else will have the money or market power to license that much material.
This doesn't even affect the major AI companies that are doing training, does it? Don't you think Google, OpenAI, Anthropic all have their own datasets by now? If they wanted to use bulk content from pirate ebook and academic paper libraries, they would've already mirrored most of those sites a long time ago.
https://web.archive.org/web/20240519104217/https://old.reddi...
magnet:?xt=urn:btih:a904e660355c49006b2e7d43893d31bf3c2be9cc&dn=libstc2.jsonl.zst&tr=udp://tracker.opentrackr.org:1337/announce&tr=https://tracker1.ctix.cn:443/announce&tr=udp://open.demonii....
> More than 400,000 fiction and non-fiction book full-texts. Multiple languages, curated, deduplicated.
> More than 6,000,000 scholarly publications, magazines, and manuals full-texts. Multiple languages, curated, deduplicated.
> 150,000,000 metadata records
There's not always going to be a correspondence between what is wrong from a legal pov and what is wrong from a moral pov.
What's this number for an LLM?
Be a shame if someone thought about it.
Edit: also, I don’t believe court decisions can be enforced retroactively so existing LLMs would be safe but I’m most definitely not a lawyer.
Books as lines of text are not human-consumable (by reasonable humans).
There's a strong case for sharing this being fair use, even if distributing hundreds of thousands of ebooks would ordinarily violate copyright.
I don't know what OpenAI's training corpus is or how they got it, but in case anyone has forgotten, Google has its own license-free collection of scanned books. It had already run that through OCR (and probably periodically re-OCRs) to generate its google books indexes, so there is no copyright argument against Google's book-trained AI models (assuming they train on books).
An argument that Books4 is piracy is an argument for an oligopoly on good AI models.
If you're using the resulting model trained with this for (academic) research (independently or at a university), that's true. If you're earning money with the model (cough chatGPT, Bard, Gemini, et. al cough), that's false.
Edit: Original version of the above paragraph started as "If you're doing...", making the below comment true. Edit is made to clarify the point further.
ETA: If you're complaining about use of the model, this becomes a dispute over how close to an original work an output can get before it becomes a copyright violation. That seems unresolvable. Look at court cases where musicians sue over short sequences of notes. Nobody knows where the threshold should be for that, for poetry, for flash fiction, for academic papers, for novels, for paintings, or for voice samples. Courts just go with whatever they think feels right in a particular case. That is an untenable legal situation in a world with AI models like these.
Also, to reiterate, Google has its own fulltext, license-free corpus of books, which it undoubtedly has used to train models. What's your view on that? Should Google be the only entity allowed to train on large books datasets because it happened to get one through a loophole of collaborating with libraries to OCR their collections?
Yeah, you're right. I fixed my comment to clarify my point, thanks.
> Also, to reiterate, Google has its own fulltext, license-free corpus of books, which...
Loopholes doesn't invalidate/override morals. I find training of any model with a corpus without consent from their respective authors immoral, regardless of laws around it.
Same for code, images, sound, text, whatever. These systems can leak their training data, or recreate them verbatim with the correct prompts. Furthermore, "claimed clean" datasets like "The Stack" are not clean by any means because the dataset is huge and tools are not good enough. Plus, even if the licenses allow this, there's still morality/consent aspect to it.
BTW, Fair Use explicitly defines the use should be non-profit. If you profit from the resulting model, it's no fair use by definition.
Any large artistic work which can be reproduced by these systems should be ingested with consent, period.
To be clear, I use none of the AI systems available on the market today.
Note the irony, too: It's the publisher writing about morality on behalf of the author. It's the publisher that would give consent for any AI-training use in your concept of the proper legal and moral order of things. The author probably doesn't know much about copyright law, but wants the work monetized as much as possible, because they like having a house and eating food, and sure, creating more content sometimes. To that end, they've assigned copyright to their publisher. Or independent creators would have to create a union and assign licensing power to the union leadership. But again, only licensing to big tech, because nobody else could afford it.
Also see the ETA: paragraph in GP, regarding the difficulty of judging whether simple works are copyrightable, or whether complex works have actually been reproduced in a copyright-violating way.
This feels like an argument over angels on the head of a pin. Look at the markets. They're all-in on AI. Big tech (and big startup) AI models will move forward, no matter what licensing agreements between publishers and AI behemoths, or copyright exceptions, end up being necessary. The only question is whether others who don't have that power or money should be able to try to train their own models (admittedly probably much lower-parameter, because they don't have datacenters full of H100s) on datasets that are at least in the same ballpark.
Some countries have a legal concept that is called that, distinct from the usual commercial rights.
Did someone hack an e-book store? Or somehow extract data from Google books?
- LibGen collections gathered over decades
- Sci-Hub
- Additional fresh scholar papers and books collected by Library STC
- bought/library borrowed books with DRM removed and shared
- scanned & OCRed books from archive.org & related projects
- non-DRM bough & shared books
Perhaps this is why so many AI books are made available for free online by their authors? For example Sutton and Barto, and Goodfellow are both available online from their authors.
Perversely I am much more likely to buy a book if the authors make it available for free. I find a hard copy more useful to study and I prefer to pick a book after sampling it (as you would in a bookshop I guess). I wouldn't go as far as getting an illegal copy to preview so those books without an easy way to review are a lot less likely to be bought (by me).
I guess if you’re a famous one then you can probably get that , but now for average people?
This is really frustrating, because as an author you can’t offer your own book on your own website for free, while on some third party websites it’s already been downloaded thousands of times.