What you're left with is a machine that produces "things that strongly resemble the original, that would not have been produced, had you not fed the original into the machine."
The fact that there's no "exact copy inside" the machine seems a lot like splitting hairs; like saying "Well, there's no paper inside the hard drive so the essence of what is copyable in a book can't be in it"
A program that always produces copies is the same as a copy. A program that merely can produce copies categorically is not.
The Library of Babel[1] can produce copyrighted works, and for that matter so can any random number generator, but in almost every normal circumstance will not. The same is true for LLMs and diffusion models. While there are some circumstance that you can produce copies of a work, in natural use that's only for things that will come up thousands of times in its training set -- by and large, famous works in the public domain, or cultural touch-stones so iconic that they're essentially genericized (one main copyrighted example are the officially released promo materials for movies).
Copyright infringement doesn't require exact copies.
Again: Imagine two AI machines, different in one way: One of them has been fed "Article X" and the other hasn't.
You press buttons on the machine(s) in the same way.
The machine that was fed "Article X" spits out something that looks like Article X, and the one that wasn't, doesn't.
The magic inside, I don't think will much matter.
Courts can call experts to testify on matters requiring specialized knowledge or expertise.
- You put thing into the machine
- You press buttons, it makes obvious derivative work
- You don't put thing into the machine, and it can't do that anymore.
There is "something" in there GENERATING COPIES and we see exactly where it came from, even if we can't identify it in the code or whatever.
If it output those reviews verbatim, sure I can see the issue, the model is over fitting. But if I tweak the model or filter the output to avoid verbatim excerpts, does an amazon lawyer have a solid footing for a "violation of copyright" lawsuit?
Honestly, instead of trying to cleanup the output, it's much safer to create a licensed input corpus. People haven't because it's expensive and time consuming. Every time I engage with an AI vendor, my first question is do you indemnify from copyright violations of your output. I was shocked that Google Gemini/Bard only added that this year.
I mean recording a good song is hard. Generating a good song almost impossible. But my gut feeling would've been that recreating a popular song for plausible deniability would be a lot easier.
Same with republishing bestselling books and related media. (I.e. take Lord of the rings and feed it paragraph for paragraph into an LLM that you've prompted to rephrase each to a currently bestselling author.)
While the extremes are obvious, there's a big stretch of gray in the middle. A similar issue occurs in non-AI art, the difference between inspiration and tracing/copying isn't well defined either, but the current method of dealing with that (being on a case-by-case basis and a human judging the difference) clearly cannot scale to the level that many people intend to use these tools.
I am simply not aware of anyone successfully doing this.
Same with the idea of "prompting" and the amount required to generate that copywritten output - again there's the extremes of "The prompt includes copywritten information" to "Vague description".
Arguably some of the same issues exist outside AI, just it's accessibility, scale, and lack of a "Legal Individual" on one side complicates things. For example, if I describe Micky Mouse sufficiently accurately to an artist they reproduce it to the degree it's considered copyright infringement, is it me or the artist that did the infringement? Then what if the artist /had/ seen the previously copywritten artwork, but still produced the same output from that same detailed prompt?
It literally says that within ChatGPT is stored, verbatim, large archives of NY Times articles and that they were able to retrieve them through their API.
Which I agree is problematic, and OpenAI doesn't have the right to disseminate that.
But that doesn't mean OpenAI doesn't have the right to train on it.
Content creators are doing a purposeful slight of hand to confabulate "outputting copyrighted data" with "training on copyrighted data".
It's illegal for me to read an NYT article and recite it from memory onto my blog.
It's not illegal for me to read an NYT article and write my own summary of the article's contents on my blog. This has been true forever and has forever been a staple in new content creation.
I'm sure if I had a JPEG of some copyrighted raw image it could still be argued that it is the same image. JPEG is imperfect, the result you get is the same every time you open it but it's not the same as the original input data.
ChatGPT would give you the same output every time, and it does if you turn off the "temperature" setting. Introduce a bit of randomness into a JPEG decoder and functionally what's the difference? A slightly different string of tokens for ChatGPT versus a slightly different collection of pixels for a JPEG.
I disagree.
If you can get the model to output an article verbatim, then that article is stored in that model.
Just because it’s not stored in the same format is meaningless. It’s the same content regardless of whether it’s stored as plaintext, compressed text, PDF, png, or weights in a model.
Just because you need an algorithm such as a specialized prompt to retrieve this memorized data, is also irrelevant. Text files need to be interpreted in order to display them meaningfully, as well.
You can't get it to do that, though.[1]
The NYT vs OpenAI case, if anything, shows that even with significant effort trying to get a model to regurgitate specific work, it cannot do it. They found articles it had overfit on due to snippets being reposted elsewhere across the internet, and they could only get it to output those snippets, and not in correct order. The NYT, knowing the correct order, re-arranged them to fit the ordering in the article.
Even doing this, they were only able to get a hundred or so words out of the 15k+ word articles.
No one who knows anything about these models disagrees that overfitting can cause this sort of behavior, but the overwhelming majority of the data in these models is not overfit and they take a lot of care to resolve the issue - overfitting isn't desirable for general purpose model performance even if you don't give a shit about copyright laws at all.
People liken it to compression, like the GP mentioned, and in some ways, it really is. But in the most real sense, even with the incredibly efficient "compression" the models do, there's simply no way for them to actually store all this training data people seem to think is hidden in there, if you just prompt it the right way. The reality is only the tiniest fraction of overfit data can be recovered this way. That doesn't mean that the overfit parts can't be copyright infringing, but that's a very separate argument than the general idea that these are constantly putting out a deluge of copyrighted material.
(None of this goes for toy models with tiny datasets, people intentionally training models to overfit on data, etc. but instead the "big" models like GPT, Claude, Llama, etc.)
1. https://fingfx.thomsonreuters.com/gfx/legaldocs/byvrkxbmgpe/...
> Even doing this, they were only able to get a hundred or so words out of the 15k+ word articles.
OK, that’s less material than I believed, which shows the details matter. But we agree that the overfit material, while limited, is stored in the model.
Of course, this can be (and surely is) mitigated by filtering the output, as long as the product is the output and not the model itself.
I disagree. Granted I'm a layman and not a lawyer so I have no clue how the court feels. But I can certainly make very specialized algorithms to produce whatever output I want from whatever input I want, and that shouldn't let me declare any input as infringing on any rights.
For the reducto ad absurdum example: I demand everyone stops using spaces, using the algorithm 'remove a space and add my copyrighted text' it produces an identical copy of my copyrighted text.
For the less absurd example.. if I took any clean model without your copyrighted text, and brute forced prompts and settings until I produced your text, is your model violating the copyright or is my inputs?
I don't think so, I think it's usually argued as two different things.
The "training on copyrighted data" argument is usually that we never licensed this work for this sort of use and it is different enough from previously licensed uses that it should be treated differently.
The "outputting copyrighted data" argument is somewhat like your output is so similar as to constitute a (at least) partial copy.
Another argument is that licensed data is whitewashed by being run through a model. So you could have GPL licensed code that is open source run through a model and then output exactly the same but because it has been outputted by the model it is considered "cleaned" from the GPL restrictions. Clearly this output should still be GPL:ed.
> It's not illegal for me to read an NYT article and write my own summary of the article's contents on my blog. This has been true forever and has forever been a staple in new content creation.
What if I compress the NYT article with gzip? What if I build a LLM model that always replies with the full article within 99% accuracy? Where is the line?
This is not a technical issue, we need to decide on this just like we did with copyright, trademarks, etc. Regardless of what you think this is not a non-issue and we cant use the same rules as we did up until now unless we treat all ML systems as either duplication or humans and neither seems to solve the issues.
I don't think anybody is making that argument. The NY Times claims to have gotten ChatGPT to spit out NY Times articles verbatim but there is considerable doubt about that. Regardless, everyone agrees that a verbatim (or close to) copy is copyright violation, even OpenAI. Every serious model has taken steps to prevent that sort of thing.
I think most people would agree that function is copyrightable if recreated verbatim.
It’s not that clear-cut. It falls into the “Fair use doctrine”The cose 107 of the US copyright law states that the resolutiodepends on>
> (1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (4) the effect of the use upon the potential market for or value of the copyrighted work.
Another thing we need to consider is that the law was redacted with the human mind limitations as a unconcious factor, (i.e not many people would be able to recite War and peace verbatim from memory). This just brings up the fact that copyright law needs a complete re-think.
This does not match my understanding of the information available in the complaint. They might claim they were able to do this, but the complaint itself provides some specific examples that OpenAI and Microsoft discuss in a motion to dismiss... and I think the motion does a very strong job of dismantling that argument based on said examples.
https://fingfx.thomsonreuters.com/gfx/legaldocs/byvrkxbmgpe/...
Would it be possible if Studio Ghibli images had not been used in the training?
otherwise those would just be unknown words, same as asking an artist to do that without any examples.
though I am curious how performance would differ between training on only actual studio Ghibli art, only fan art, or a mix. Maybe the fan art could convey what we expect 'studio Ghibli style' to be even more, whereas actual art from them could have other similarities that that tag conveys.
If the produced work is not a copy, why does it matter if it was generated by a biological brain or by a mechanical one?