When model training reads the text and creates weights internally, is that a substantial transformation? I think there’s a pretty strong argument that it is.
When model training reads the text and creates weights internally, is that a substantial transformation? I think there’s a pretty strong argument that it is.
The point here is that book files have to be copied before they can be used for training. Copyright texts typically say something like "No unauthorised copying or transmission in any form (physical, electronic, etc.)"
Individuals who torrented music and video files have been bankrupted for doing exactly this.
The same laws should apply when a corporation downloads torrent files. What happens to them after they're downloaded is irrelevant to the argument.
If this is enforced (still to be seen...) it would be financially catastrophic for Meta, because there are set damages for works that have been registered for copyright protection - which most trad-pubbed books, and many self-pubbed books, are.
Only if they seeded the data and some other entity downloaded it, i.e. they hosted the data. In a previous article I believe it was called out that Meta was being a leecher (not seeding back what they downloaded).
It's the hosting that gets you, not the act of downloading it.
The lender owns the book, and it is within his rights to loan it to whoever he wants. That is legal. Making this illegal would end libraries.
The borrower is well within his rights to accept the book, and as the current owner he is even allowed to make a copy of the book (see the famous TIVO case). Making this illegal would end backups and format/time shifting.
When the borrower returns the book, he keeps the copy. Oh no! Surely he must now become a criminal? Nope. Possessing an unauthorized copy is also not illegal, despite what many copyright holders would like you to believe. Making this illegal would also criminalize a lot of legitimate format/time shifting, again see the famous TIVO case.
If the borrower were to loan his homemade copy to someone else THEN it would finally become illegal.
Nothing about AI changes any of this.
But if you make books contents available online via some service that regurgitates its contents you would be totally sus because you can be considered in a business of selling derivative works.
I download a torrent with movie that I didn't pay for. If I don't allow to seed it, then I don't get in trouble. If I let it seed either during the download process or after, I'd get a DMCA notice if that torrent/magnet link was getting tracked.
I don't need a hypothetical book, that is just how it works if I were to download illegally obtained documents/media.
As technical as people are in this thread, easy to tell when folks didn't have their parents wondering why they were getting scary letters from the ISP.
In Googles case they were digitizing the books (that they did not own), and publishing snippets for search users to help them find books and other material that weren't indexed on the web. The court found they had that right, but did place some pretty strict limits on them.
Still, Google was allowed to keep their database of scanned material despite not owning the originals.
Link: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
Someone photocopying a book to read on the toilet (and leave the original on their nightstand) isn't engaging in transformative use. They're also harming the market for the work because if they hadn't made this photocopy, they would've had to buy a second copy of the book to get the same benefit.
The situation you describe is more akin to "format shifting" or maybe "space shifting", which is converting copyrighted material to other formats or places as a backup, for preservation purposes, or just for convenience. It is legally protected in the US, and most of the rest of the world. I believe that was settled due to early litigation with VCR's, but there are mountains of relevant cases at this point.
I do recall reading that the UK recently changed their laws in that regard, and like most copyright law it can depend on the specifics of the situation and even (*yuck*) intent. So it's always worth checking if you're unsure.
Citation needed.
However, people have been prosecuted for not even hosting a torrent, but merely providing a link to where people can find it.
e.g. https://torrentfreak.com/operator-of-popcorn-time-info-site-...
Seems like a big gap there.
The computer model is working differently of course but functionally it's the same idea.
How is this done? Are bits not written into RAM or disk? Are they not sent between machines in a training cluster? That's copying.
> it is seemingly not far removed from how humans consume content
Except that humans don't make full copies to RAM, or disk or paper.
AI doesn't need lasting copies to train, however I don't know what the actual implementation is. But if it's ruled that they can only use copyrighted data if it's not stored for more than the time it would take a human to consume, It wouldn't really cripple the models, but perhaps make training more logistically challenging.
It's important to understand that models are not data archives. They are statistical constructs made from getting quizzed, that uses human made content to generate the quiz questions.
Wired explicitly sent that article to their computer for the purposes of reading it so it's not a copyright violation.
Images on your retina form exact copies.
They are scanned and translated into impulses that are then sent to a first set of "neural columns" - that's an exact copy.
This is then connected to the visual cortex by the two most high bandwidth links in the human body ("the optical nerve", there's 2 of them of course, always wondered why everybody insists on using the singular). Why would you have that high bandwidth link unless to create verbatim copies.
The way those columns are structured also very strongly suggests they make carbon copies, which they then make available on the "brain bridge" (which is probably at least vaguely similar to the "attention matrix" of a transformer). If it does work like that, that's also a verbatim copy.
The only way "humans don't make full copies to RAM" is that humans don't have separate RAM. The processing power is colocated with the processing, even on a microscopic level. You know, what everybody knows is the best way of doing things even in silicon, it's just incredibly impractical if you can't rebuild your circuit every time there's a slight change to the instructions your "computer" carries out (the brain is not a "Von Neumann architecture", except it kind of is when it regrows connections. But in the short term it isn't)
Not for the purposes of copyright law.
> is that humans don't have separate RAM [or disk]
And that turns out to be incredibly important. Humans can't create a lasting, shareable copy of a copyrighted work by consuming it.
And that's a copyright violation.
> Using software almost always involves creating copies, even though many of these copies only exist for a very short time. For example, executing a program means copying it from the hard disk into RAM so that the CPU can interpret the instructions. Because of this, the right to run a program is considered to fall under the copyright of the author.
For comparison, when a human looks at the letters, there is no copying.
Also, models can reproduce text verbatim which proves that they store it.
So it is unfair when ordinary folks got sued for this and Zuckerberg wants to get away with a million times larger violation. He must go directly to jail.