Another way of thinking of it is compressed knowledge, with learned structure, and an ability to pick likely words to describe that structure rather than an ability to use original words.
The compression ratio of the source corpus versus the model demonstrates it doesn't have all the source material, ipso facto, results are not infringement.
As an easy to understand analog, if you grab a Shutterstock photo, change it from 100% quality to 5% quality, the file size drops to 100th of the original and on display, none of the pixels are the same. While it may be the same general concept depicted badly, it is certainly not the original and clearly isn't copyrightable.
(Note: That last paragraph is firmly tongue in cheek /s ... I've used the "uniqueness (originality) is compressed out, what remains as learned is fair use" argument in the past, but on further reflection about the interaction of creativity and structure, I'm not sure it stands, especially at temperature 0 aka zero creativity.)
If a human didn't create it, it can't have a copyright.
I draw an image. I write some lines of code that draws an image. I write many lines of code that do something, gather something, and then draws an image?
I just wonder how one can reasonably draw hard lines (and now I don't mean art :D) in the future.. there are so many ridiculous incarnations of copyright and patent fights already out there..
In the context of using AI or coding, the question is how much creative control you have over the output. When writing code to draw an image pixel by pixel, then you have full control, probably even for procedural art when you devised the algorithm completely on your own. When using an existing library that can draw crayon strokes or produce nice water color fills, there is less attributable to your own creativity. Similar for AI, where it will have to be assessed how much of the creativity is attributable to the human-invented prompting (likely not much) vs. to the AI.
There are lots of instances of fairly fragile architectures that work really well in a particular application, but if you change some small part of that architecture, they fail. Now say I spend a lot of research time identifying that architecture and maybe even patent it as a process. I could even publish the methods of exactly what that process entails and somebody could copy it an use it. But I think that would be illegal (assuming AI can be patented in this way).
Training an ML model doesn't seem to fall under the established exceptions which qualify as "fair use".
Given most of what we are talking about would fall under the output being used commercially, lets look at the end of of fair use test.
> the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
>the effect of the use upon the potential market for or value of the copyrighted work.
Unless your opponent is using a model exclusively trained on Twilight to output a Twilight sequel before you the author, you are going to have a have a hard time saying they used a large part of your work and that it directly displaces a sale of your work.
A collage of word samples from millions of works is going to fall under fair use. If the model did a good enough job of compressing your work in its entirety, and it outputs that copy, you might have a case.
> The compression ratio of the source corpus versus the model demonstrates it doesn't have all the source material, ipso facto, results are not infringement.
What?! This is not at all how copyright law works and that's also not at all how LLMs or compression work! You're wrong about both the law and the technology!
Just because the model can't losslessly reproduce the entire training dataset doesn't mean that the model can't losslessly reproduce copyrightable fragments of the training dataset.
And the law truly doesn't give a damn that a verbatim copy of a piece of code or sequence of paragraphs came out of a model instead of a copy/paste from github. If it's the same basket of bits, or close enough in certain cases, it's covered by copyright, full stop.
> As an easy to understand analog, if you grab a Shutterstock photo, change it from 100% quality to 5% quality, the file size drops to 100th of the original and on display, none of the pixels are the same. While it may be the same general concept depicted badly, it is certainly not the original and clearly isn't copyrightable.
As an easier to understand analog, if you grab a BITMAP of a Shutterstock photo and convert the file type to a reasonable JPEG (lossy compression!!!), tht's very clearly still covered by the copyright on the .bmp.
Firmly tongue in cheek sarcasm, as in, this is obviously wrong.
Or, why making and redistributing a JPEG of a copyright-protected artwork is not infringement.
Results decidedly not guaranteed if you try this “a copy with lossy compression is not infringement” argument in an actual court of law.
You've replied in the spirit intended. :-)
Let's take that to the next step - what happens if ChatGPT happens to get trained on something that is copyrightable - then a prompt script can be used to reproduce the original - does that mean the original is now considered not protected by copyright?
This is a lot more murky than "show me the prompts/parameters" and it's not protected by copyright.
If a copyright precedes the existence of a model, the model was trained on that code and someone distributes code from the model that is a near-verbatim reproduction, that's a different story (in large part because getting that near verbatim reproduction would probably require willful intent rather than just being accidental)
Just out of curiosity.
I don’t think the owner of the copyright loses copyright simply because something else output a copy of his work. Maybe it is “generated” but the copyright holder still has his rights per copyright statutes and you would still be violating those rights by distributing your own copies from the AI output.
OpenAI ToS 3(a):
"OpenAI hereby assigns to you all its right, title and interest in and to Output. This means you can use Content for any purpose, including commercial purposes such as sale or publication, if you comply with these Terms."