If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement?
How many times does one need to compress the JPEG before it's fair use? I'm legitimately curious what the test is here.
The model isn’t storing the book.
I’m sure all these ‘clever’ questions would be useful if this trial was about humans but it’s not.
The training material is used to play this guessing game to dial in it's weights. The training data is picked up, used as reference material for the game, and then discarded. It's hard to place this far from what humans do when reading, because both are using the information to mold their respective "brains" and both are doing an acquire, analyze, discard process.
At no point is training data actually copied into the model itself, it's just run past the "eyes" of the model to play the training game.
I think that is the center of the conversation. What does it mean for a computer to "understand"? If I wrote some code that somehow transformed the text of the book and never return the verbatim text but somehow modified the output, I would likely not be spared because the ruling will likely be my transformation is "trivial".
Personally, I think we have several fixes we need to make:
1. Abolish the CFAA.
2. Limit copyright to a maximum of 5 years from date of production with no extension possible for any reason.
3. Allow explicit carveout in copyright for transformational work. Explicitly allow format shifting, time shifting, yada yada.
4. Prohibit authors and publishers from including the now obviously false statements like "No part of this publication may be reproduced, distributed, or transmitted in any form or by any means, including photocopying, recording" bla bla bla in their works.
5. I am sure I am missing some stuff here.
For brand protection, we already have trademark law. Most readers here already know this but We really should severe the artificial ties we have created between patents, trademarks, and copyright.In early computing, everything was closed sourced. Quoting the wikiepdia page,
To develop a legal BIOS, Phoenix used a clean room design. Engineers read the BIOS source listings in the IBM PC Technical Reference Manual. They wrote technical specifications for the BIOS APIs for a single, separate engineer—one with experience programming the Texas Instruments TMS9900, not the Intel 8088 or 8086—who had not been exposed to IBM BIOS source code.
The legal team at Phoenix deemed inappropriate to "recall source in their own words" for legal reasons.
My non-legal intuition is that these companies training their models are violating copyright. But, the stakes are too high--it's too big to fail if you will. If we don't do it, then our competitors will destroy us. How do you reconcile that?
But any such arrangement needs to be hammered out by the legislature. As laws are, I think it's pretty clear that infringement is happening.
When model training reads the text and creates weights internally, is that a substantial transformation? I think there’s a pretty strong argument that it is.
The computer model is working differently of course but functionally it's the same idea.
How is this done? Are bits not written into RAM or disk? Are they not sent between machines in a training cluster? That's copying.
> it is seemingly not far removed from how humans consume content
Except that humans don't make full copies to RAM, or disk or paper.
AI doesn't need lasting copies to train, however I don't know what the actual implementation is. But if it's ruled that they can only use copyrighted data if it's not stored for more than the time it would take a human to consume, It wouldn't really cripple the models, but perhaps make training more logistically challenging.
It's important to understand that models are not data archives. They are statistical constructs made from getting quizzed, that uses human made content to generate the quiz questions.
Wired explicitly sent that article to their computer for the purposes of reading it so it's not a copyright violation.
Images on your retina form exact copies.
They are scanned and translated into impulses that are then sent to a first set of "neural columns" - that's an exact copy.
This is then connected to the visual cortex by the two most high bandwidth links in the human body ("the optical nerve", there's 2 of them of course, always wondered why everybody insists on using the singular). Why would you have that high bandwidth link unless to create verbatim copies.
The way those columns are structured also very strongly suggests they make carbon copies, which they then make available on the "brain bridge" (which is probably at least vaguely similar to the "attention matrix" of a transformer). If it does work like that, that's also a verbatim copy.
The only way "humans don't make full copies to RAM" is that humans don't have separate RAM. The processing power is colocated with the processing, even on a microscopic level. You know, what everybody knows is the best way of doing things even in silicon, it's just incredibly impractical if you can't rebuild your circuit every time there's a slight change to the instructions your "computer" carries out (the brain is not a "Von Neumann architecture", except it kind of is when it regrows connections. But in the short term it isn't)
Not for the purposes of copyright law.
> is that humans don't have separate RAM [or disk]
And that turns out to be incredibly important. Humans can't create a lasting, shareable copy of a copyrighted work by consuming it.
And that's a copyright violation.
> Using software almost always involves creating copies, even though many of these copies only exist for a very short time. For example, executing a program means copying it from the hard disk into RAM so that the CPU can interpret the instructions. Because of this, the right to run a program is considered to fall under the copyright of the author.
For comparison, when a human looks at the letters, there is no copying.
Also, models can reproduce text verbatim which proves that they store it.
So it is unfair when ordinary folks got sued for this and Zuckerberg wants to get away with a million times larger violation. He must go directly to jail.
The point here is that book files have to be copied before they can be used for training. Copyright texts typically say something like "No unauthorised copying or transmission in any form (physical, electronic, etc.)"
Individuals who torrented music and video files have been bankrupted for doing exactly this.
The same laws should apply when a corporation downloads torrent files. What happens to them after they're downloaded is irrelevant to the argument.
If this is enforced (still to be seen...) it would be financially catastrophic for Meta, because there are set damages for works that have been registered for copyright protection - which most trad-pubbed books, and many self-pubbed books, are.
Only if they seeded the data and some other entity downloaded it, i.e. they hosted the data. In a previous article I believe it was called out that Meta was being a leecher (not seeding back what they downloaded).
It's the hosting that gets you, not the act of downloading it.
The lender owns the book, and it is within his rights to loan it to whoever he wants. That is legal. Making this illegal would end libraries.
The borrower is well within his rights to accept the book, and as the current owner he is even allowed to make a copy of the book (see the famous TIVO case). Making this illegal would end backups and format/time shifting.
When the borrower returns the book, he keeps the copy. Oh no! Surely he must now become a criminal? Nope. Possessing an unauthorized copy is also not illegal, despite what many copyright holders would like you to believe. Making this illegal would also criminalize a lot of legitimate format/time shifting, again see the famous TIVO case.
If the borrower were to loan his homemade copy to someone else THEN it would finally become illegal.
Nothing about AI changes any of this.
But if you make books contents available online via some service that regurgitates its contents you would be totally sus because you can be considered in a business of selling derivative works.
I download a torrent with movie that I didn't pay for. If I don't allow to seed it, then I don't get in trouble. If I let it seed either during the download process or after, I'd get a DMCA notice if that torrent/magnet link was getting tracked.
I don't need a hypothetical book, that is just how it works if I were to download illegally obtained documents/media.
As technical as people are in this thread, easy to tell when folks didn't have their parents wondering why they were getting scary letters from the ISP.
In Googles case they were digitizing the books (that they did not own), and publishing snippets for search users to help them find books and other material that weren't indexed on the web. The court found they had that right, but did place some pretty strict limits on them.
Still, Google was allowed to keep their database of scanned material despite not owning the originals.
Link: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
Someone photocopying a book to read on the toilet (and leave the original on their nightstand) isn't engaging in transformative use. They're also harming the market for the work because if they hadn't made this photocopy, they would've had to buy a second copy of the book to get the same benefit.
The situation you describe is more akin to "format shifting" or maybe "space shifting", which is converting copyrighted material to other formats or places as a backup, for preservation purposes, or just for convenience. It is legally protected in the US, and most of the rest of the world. I believe that was settled due to early litigation with VCR's, but there are mountains of relevant cases at this point.
I do recall reading that the UK recently changed their laws in that regard, and like most copyright law it can depend on the specifics of the situation and even (*yuck*) intent. So it's always worth checking if you're unsure.
Citation needed.
However, people have been prosecuted for not even hosting a torrent, but merely providing a link to where people can find it.
e.g. https://torrentfreak.com/operator-of-popcorn-time-info-site-...
Seems like a big gap there.
It seems like it is very much a matter of fidelity.
As mentioned in another comment, LLMs (and most popular machine learning algorithms) can be viewed, correctly, as compression algorithms which leverage lossy encoding + interpolation to force a kind of generalization.
Your argument is that a video wouldn't count as pirated if the compression used for the pirated copy was lossy (or at least sufficiently lossy). The closest real world example would be the cases where someone records a the filming of a movie on their phone then uploads it. Such a copy is lossy enough that you can't produce anything really like the original, but my most definitions is still considered copyright.
You would never use a human to backup your financial reports, but the human might be able to give a good overview. You would never use an LLM to backup your financial reports, but they might be able to give a good overview.
AI training data is disposable. There is nothing that could be called a compression algorithm that disposes all of the data you put into it. AI uses training data as examples of what the next token in a token sequence is. The examples are disposable reference points, not the model itself. That's how you get image models that are 20GB in size despite training on 20PB of data. It's 20PB of examples used to form the shape of a 20GB model. You could show it 5GB of training data or 500EB of training data and it would still be 20GB - because it is not a compression algo, it's a 20GB shape formed by external data.
I'm sorry, but this a fundamentally incorrect view of machine learning (including, but not limited to transformers).
From an information theoretic perspective the two are essentially identical with the exception that standard compression algorithms do not have a proper "loss" function other than just trying to minimize reconstruction loss with the resulting compression size.
Here's a link to the section on the Wikipedia for more information if you'd like [0]. MacKay's Information Theory, Inference and Learning Algorithms is the standard full text treatment of this topic [1]. Ted Chiang's article "ChatGPT is a Blurry JPEG of the web" is pretty good "pop sci" exploration of this topic if you don't want to get too into the mathematics [2].
0. https://en.wikipedia.org/wiki/Data_compression#Machine_learn...
1. https://www.inference.org.uk/itprnn/book.pdf
2. https://www.newyorker.com/tech/annals-of-technology/chatgpt-...
Humans are totally capable of data compression. This will just devolved into a semantics game of what a data compressor is.
LLMs were not developed to be, do not function as, and are not use as data compression utilities. Please, come knocking when a service provider exists that will use LLM's to compactly store your company data.
Again, from a information theoretic view point, this is exactly what they are doing, how they where developed and how they function.
I don't know any serious researcher in ML that would find this claim even remotely controversial. It's really not just "a semantics game", its a part of a foundational understanding of the topic. If you want to understand LLMs from this perspective, a good place to start is with an auto-encoder which does try to learn a standard compression algorithm, the move on to more sophisticated embedding models (found in a lot of recommender systems) which try to learn an additional objective on top of minimizing reconstruction error. You'll then see that Transformers and all other major NN architectures fall out of these basic principles.
> Please, come knocking when a service provider exists that will use LLM's to compactly store your company data.
This is literally what every vectordb company does right now, as well as all "chat with your docs" type startups.
Can you get transformers to regurgitate information verbatim? Yes.
Would anyone in their right mind rely on a transformer to do so? No.
Would anyone in their right mind rely on a vectorDB to do so? No.
Would anyone in their right mind use a vectorDB/RAG/SQL/transformer combo to do so? Yes.
Is youtube going to drop VP9 for GeminiEncode to save google billions in bandwidth? No.
https://arxiv.org/pdf/2309.10668
Transformers are also used in the top algorithm right now on the Large Text Compression Benchmark. https://bellard.org/nncp/nncp.pdf
You can compress 20PB of text to 20Gb or even less, if input is super repetitive. So the same with images, if 50% of the images are cats then you learn how to represent the cat pixels with a few vectors and then you could represent all the cats int he world doing all possible cat actions.
But please have the courage to respond to this, when the AI is caught regurgitating the exact text from a popular book, the exact verses from a poem, the exact code function from some code , then how can you defend that is not memorizing things? If a human uses my poem(after they read it) and signs his name under it would you defend them?
And yes LLMs can recall exact material, but it is excerpts and fragments. There is statistical significance to it's ordering. Humans readily do this too (excerpts and fragments), most artists can draw a batman symbol (but not an episode of batman). That doesn't in anyway mean that artists should not be allowed to ever see a batman symbol. It means that artists shouldn't be allowed to get paid to draw one. And they are not. And LLMs are not exempt either.
But the fix is output filtering, just like everything else that can violate copyright. Which is already being done (albeit poorly, but way better than 2 years ago), the same as artists will not draw the batman symbol for you despite being able to.
maybe even simpler, I create a zip format where I randomly replace words with their synonym, or group of words with something equivalent. Would you defend this as original ? why my random transformations are not original while mathematical transformations you will defend ?
And how can you suggest putting output filters to protect only the giants for copyright and everyone else gets screwed.
We should be building robust copyright filters and everyone should be able to contribute their work to it.
but that is a different issue than whether or not an LLM is legally allowed to view a work that is publicly available.
Again, pretty much every artist is capable of off-hand copyright violation on the spot. This has been true forever. We don't bar them from seeing art to prevent this.
We are not talking about the scenario where LLM does a web search and "transforms" your blog as a response to a user prompt.
We are talking about people writing code to torrent content, including copyrighted content as books and newspaper articles then this same people create soem software/script/blackbox and they put this copyrighted stuff inside and this same people also put a filter in the output to attempt and block SOME of the copyrighted material to be spitted out(they did it for Dune, Harry Poter , probably other popular stuff but for sure they did not done it for Romanian copyrighted material or some small writer).
If I publish a poem I do not give implicit permission to anyone to use it as they see fit, and have scripts/software/backboxes train on it and spit parts of my poem out.
Needing the original material isn't enough for claiming copyright infringement as we have existing counter examples
Nobody is going to try to extract pages by pages a book from ChapGPT, let's be realistic. (and you can't anyways)
I don't think the current state of LLM would be able to write 200 pages in a coherent manner anyways.
That is why so called derivative works are allowed (and even encouraged). If copyrighted material is ingested, modified or enhanced to add value, and then regurgitated that is legal, whereas copying it without adding value is not legal.
If derivative works weren't deemed acceptable copyright would have the opposite of it's intended effect and become an impediment to progress.
Derivative works are tolerated in some cases like some manga or fanfics but it is a gray area and whenever the author or publisher wants to pursue it it is their full right to do it. Many do pursue it
(You can get inspired by something, and this is where some arguments can happen if you get inspired too mmmm literally, but no one will say with a straight face that inspiration is a thing that happens to software)
If the work is "derivative" in the legal sense it is copyrighted, and you may not create derivative works without the copyright holders permission.
What I should have said is that simply being inspired by a work or copying unprotectable elements (like facts or ideas) does not create a derivative work.
For example, if ChatGPT were to generate Star Wars, except with Dookies instead of Wookies, that might be illegal. If it were to learn what a spaceship is from Star Wars and then create something substantially new it would not. The key is is that it must not be substantially similar to the original. You must add enough value that it becomes something new, not just rehash the original.
So… it’s complicated. This is one of the weird areas where music copyright and other copyright seem to differ in the US.
In the US the situation is complex and there are a lot of weird special interests [0], but generally a composer/author of a song has the right to decide who first records and releases the song, but after the first recording covers require a mechanical license, which is compulsory (ie: the author cannot object).
In music there are _a lot_ of special cases and different rights are decided with different kinds of licenses, some of which are compulsory. I think it’s an area that doesn’t make for good analogies with copyright in other media.
Which is compulsory for the performer too.
A derivative work like cover is sort of acceptable when it's performed by a person live for some audience (grey area but twitch sort of allows it. with a bunch of rules). As soon as you want to publish it you MUST have a license. And chatbot is a derivative work totally not performed live by a person for some audience
I saw great tracks that got taken down from all legal channels because they featured a sample from another song. Sometimes they remained up but mostly they were taken down. It is fully original publisher's discretion...
Derivative works are not given a free pass from the normal constraints of copyright. You cannot legally publish books in the universe of A Song of Ice and Fire without permission from the author (and often publisher), calling them “derivative works.”
It’s why fan fiction is such a gray area for copyright and why some publishers have historically squashed it hard.
The exceptions for this are typically fair use, which requires multi-factor analysis by the judiciary and is typically decided on a case-by-case basis.
Derivative works are not "allowed (and even encouraged)" without a license from the copyright holder. Creating a derivative work is an exclusive right of the copyright holder just like making verbatim copies and requires a license for anyone else, unless an exception to copyright protection (like fair use) applies.
That seems to go against the notion that copyright can last beyond the author's lifetime - most arts and science progress tends to reduce after death
The test is if a judge says it is fair use, nothing else.
The judge will take into account the human factor in this matter, e.g. things like who did the actual work, and who just used an algorithm (which is not the hard part anymore, code can be obtained on the internet for free). And we all know that DL is nowhere without huge amounts data.
The NYTimes in 2023 was able to demonstrate that the models can reproduce entire articles verbatim[0] with minimal coercion.
[0]https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...
It would be a remarkable quirk of statistics that, if given all text on the internet except for the NYTimes back catalogue, a model would produce any NYT article.
It becomes illegal if I try to distribute those copies
So the question is, does distributing an AI that has been trained on Harry Potter count as distributing Harry Potter?
Not of physical media. You're allowed to make archival copies of digital media.
> Or reading a book via a computer would be illegal
No you purchased a license (or your library did, in the case of e-borrowing) to read the book on a computer. That makes it legal.
This scenario seems quite contrived but is there an actual court precedent allowing it? I'm 100% confident no one will ever prosecute you for doing it but that's not the same thing as "allowed".
In another thread I already posted about https://en.wikipedia.org/wiki/American_Broadcasting_Cos.,_In....
This case was about pointing <recording equipment> at <content someone is allowed to access> and making a copy to transmit it to <only that person>. The Supreme Court held it to be illegal. There was a lot of money on the line, which is why it went so far.
Your example involves transmission and mine doesn't, and that's a whole different can of worms.
Also the result of that case was self-contradicting so it's not a great basis to build too much logic upon.
I'm not aware of this principle. Where is it spelled out?
> Also the result of that case was self-contradicting
I agree the verdict was a travesty. An innovative business went to ridiculous lengths to stay on the right side of the copyright mafia (data centers with tiny individual TV antennas for each subscriber FFS!) while providing a better product and experience. They still had the hammer brought down on them.
Well, do CDs give you a license agreement that allows you to copy the data? I've never seen one. And it's nearly impossible to play a CD without that copying.
Exactly. It's how you are supposed to use the CD. That's not true for your example of a book on a webcam. You're supposed to read the book, not an image of the book.
That's not the same thing as copying for "processing".
Would it be a violation to play back a record like a CD and have a digital buffer? That would be pretty silly.
Selling you the CD explicitly grants you the right to play it using a CD player.
> Would it be a violation to play back a record like a CD and have a digital buffer?
I genuinely have no idea lol.
While a tool being used to create infringing copies of some other work (whether or not it is the source material used to create the tool, and whether or not the infringing material is also verbatim copies) is relevant to whether the tool vendor is liable for contributory infringement for the infringing use of the tool, the absence of a capacity for creating such copies isn't usually enough to say that copying to make the tool isn't infringing.
(That said, generative AI tools, including LLMs specifically, have been shown to have the capacity to make such copies, to the extent that vendors of hosted models are now putting additional checks on output to try to mitigate the frequency with which verbatim copies of substantial portions of training-set works are produced, so arguing that LLMs can't do that is silly.)
Exactly. I asked my Gemma how long of a quote it could give me of a given book, if I were the author & gave express permission, and I was a bit surprised it readily admitted it could
> Without Permission (Current Limit): Single sentence.
> With Broad Permission (Full Reproduction Allowed): I could theoretically quote the entire book.
Eye-opening (for me, at least).
Transformers are fundamentally large compression algorithms where the target of compression is not just to minimize reconstruction loss + compressed file size. In fact, basically all of machine learning used today can be viewed through the lens of learning a compression algorithm with added goals other than the usual.
By this logic if I create a lossy Jpeg of a copyrighted image it's not "copying" because the lossy compression.
There’s plenty of case law there…
Using content to train an LLM is not copying the content. I'm ignoring the silly "but actually" arguments about the content being in RAM so it's "copying". It's using the content to generate a statistical model of token (word-ish) relationships and probabilities. If you write content that is so original in it's wording and I train an LLM against it, then there is certainly the possibility that the LLM could be provoked the recall the exact words you used. You'd have to set the parameters just right to make it happen and I think that proper training would drastically lower if not remove that possible scenario. But even if it doesn't, the LLM doesn't have a copy of that original content. All it has is weights representing those relationship probabilities. Yes, the minutia is more complex, but that is the essence. If my LLM were to generate enough of this essentially verbatim unique content and I tried to publish or copyright it, then I as the user should be on the hook. But then you get into a discussion about how many words in a unique sequence does it take to be infringement?
Obviously, I am not a lawyer.
My summation in all of this is that new laws need to be put into place to handle this stuff because the existing ones are sufficiently non-definitive and/or ill-suited such that every party is forming strong opinions about how old laws apply to new situations and causing massive friction.
If we're going that way, let me torrent every movie and TV show ever to "train" myself.
https://en.wikipedia.org/wiki/American_Broadcasting_Cos.,_In....
I'm not a legal expert. My layman's understanding of the case above is Aereo was in violation because they made copies of content - content that the receiver was already allowed to access - available over the Internet to the intended receiver. That is to say, the copying was the problem.
"Aereo's retransmission of television broadcasts was a "public performance" of the networks' copyrighted work. The Copyright Act of 1976 forbids such performances without the permission of the holder of the copyright. Second Circuit Court of Appeals reversed. Court membership"
Exactly this. Legal copying requires a license.
Where your analogy goes wrong is you're saying you want to "[Circumvent] payment to obtain copyright material for training" to use Workaccount2's words.
Because I'm certainly not allowed to photocopy a library book in its entirety. And I guarantee you a Netflix subscription doesn't allow me to keep a copy of a movie on my hard drive and use it for training man or machine.
IANAL but that probably falls under fair use? You'll get in trouble if you photocopy the work and sell access to it.
I've not found case law for that. I've had this same argument on HN multiple times over the past few months.
Copyright doesn't protect against all forms of duplication. For instance, you own the copyright to your post and grant HN a license to offer copies of it. I have no direct license from you to copy the content of your post; but I can copy it to memory, copy a cache to disk, and make a copy appear on my display.
It’s not a good example, because if you grant a license you give them the right to make copies. The problem is not when Meta got licenses, it’s when they did not.