The Authors Whose Pirated Books Are Powering Generative AI
theatlantic.com
theatlantic.com
[1] https://torrentfreak.com/anti-piracy-group-takes-prominent-a...
[2]: https://academictorrents.com/details/0d366035664fdf51cfbe9f7...
But, BitTorrent is probably the only way a dataset like this can survive now. It’s a damn shame since OpenAI are the ones making money, and researchers are simply trying to replicate their scientific efforts.
Original books3 announcement thread: https://x.com/theshawwn/status/1320282149329784833?s=61&t=jQ...
An article on books3 from Gizmodo, which I quite like: https://gizmodo.com/anti-piracy-group-takes-ai-training-data...
By the way, you can use aria2c to download just the books3 part of that torrent. Use aria2c —-select-file=44 and pass the torrent url. Takes about 30 minutes. Plus most people probably don’t have 800GB free.
They should make space. Data sets in general are hugely slept on among hobbyists.
Even just what's publicly available, like Wikipedia's dumps; you can do so much fun stuff with them and a few days of compute time!
And even the output has been ruled not copyrightable in recent court proceedings.
Edit: As a secondary point, it's not like Meta bought all those books. They downloaded unlicensed copies.
You can use books without buying them. e.g. libraries, or borrowing books from anyone even if they're not a "library".
Copyright is not about internal use, it's about copying and distribution.
This is not about what anyone thinks should be legal, but rather what is legal under the current law. The law was not designed for the digital era where "use" could be something other than a person consuming the content, and this has not been meaningfully addressed, therefore other uses are legal by default because that's how the law works.
How did they scan them in then?
By your logic, distributing the digits of pi could be construed as copyright infringement otherwise.
of course you can - pi contains all known combinations of digits.
On the other hand, the "address" of a text within a memorizing large language model would just be the prompt "give me the text of X".
The ease with which these "retrieval" operations can be done is irrelevant.
they consented by virtue of publishing the work. If a human eye can read it, then consent has been given.
and there you've just put your own words into my post. Nowhere did the idea of selling comes into this (nor distribution of any kind).
If you publish, you implicitly allow people to purchase and read your book. The information and ideas learnt off that book is not subject to copyright.
No one is asserting that computers themselves should have rights. But their users certainly do. If it is legal for me, a human person, to do something inside my own brain, why would it not also be legal for me to do a rough approximation of the same thing inside my computer?
what does 'important' mean here?
If the machine can reproduce the process of learning, there should not be a law to prevent such learning from happening. Just like there wasn't a law to prevent the looms from operating while there are weavers out there.
That isn't to say that the computer knows what it is doing and it isn't to say that the computer isn't going to produce a lot of nonsense. But even random markov chain babbling will sometimes produce interesting sentences that make you stop and think. The meaning is in the mind of the reader.
If we think of LLMs as a large mixing pot that combines ideas from vast numbers of written sources in what is often mundane, but may occasionally spark a thought or an imaginative creative leap that humans haven't yet made, they become just another tool that humans can use to think and sometimes see things differently.
I don't think authors should get royalties on GPT output, and I also don't think OpenAI needs special licensing deals in order to train, but I get that they should at least buy the books.
> For example, when an AI technology receives solely a prompt from a human and produces complex written, visual, or musical works in response, the “traditional elements of authorship” are determined and executed by the technology—not the human user. Based on the Office's understanding of the generative AI technologies currently available, users do not exercise ultimate creative control over how such systems interpret prompts and generate material. Instead, these prompts function more like instructions to a commissioned artist—they identify what the prompter wishes to have depicted, but the machine determines how those instructions are implemented in its output.
It's like pissing in the ocean and thinking you caused a tsunami in Southeast Asia.