1) "Dear artists, the model cannot infringe upon your copyright because it's merely learning like a human does. If it accidentally outputs parts of your book, you know, it just accidentally plagiarized. We all do it haha! Our attorneys remind you that plagiarism is not illegal in the US."
2) "Dear engineers, the output of our model is copyrighted and thus if you use it to train your own model, we own it."
I am not sure how both of those can be true at the same time.
Oops.
I added the "haha" in there because the probability of a human doing this kind of goes way down as the length of the text increases. Can you type, verbatim, an entire chapter of a book? I can't. But, I bet the AI can be convinced in rare cases to do that.
The whole thing is very interesting to me. There was an article on here a couple days ago about using gzip as a language model. Of course, gzipping a book doesn't remove the copyright. So how low does the probability of outputting the input verbatim have to be before copyright is lost?
Reading the book and benefitting from what you learned? Obviously not copyright infringement. Putting the book into gzip and sending your friend the result? Obviously copyright infringement. Now we're in the grey area and ... nobody knows what the law is, or honestly, even how to reason about what the law wants here. Fun times.
(Personally, I lean towards "not copyright infringement", but I'm not a big believer in copyright myself. In the case of AI training, it just makes it impossible for small actors to compete. Google can just buy a license from every book distributor. SmolStartup can't. So if we want to make AI that is only for the rich and powerful, copyright is the perfect tool to enable that. I don't think we want that, though.
My take is that the rest of society kind of hates Tech right now ("I don't really like my Facebook friends, so someone should take away Mark Zuckerberg's money."), so it's likely that protectionist laws will soon be created that ruin it for everyone. The net effect of that is that Europe and the US will simply flat-out lose to China, which doesn't care about IP.)
Of course, the models that are developed for internal use by the Chinese government won't be so limited, regardless of what the law says. But then neither be the ones developed by Western three-letter agencies. So don't worry about the "Great Game"; they'll do just fine one-upping each other and screwing over all of us in the process.
This is why fundamental "interpolative" techniques like ChatGPT (whose weights are in theory frozen) is still basically super-intelligent.
It's a well-understood concept that our minds function by making sense of the world through patterns. This is the essence of interpolation - taking two known points and making an educated guess about what lies in between. Ever caught yourself finishing someone's sentence in your mind before they do? That's your brain extrapolating based on previous patterns of speech and context. These processes are at the heart of human creativity.
The field of Cognitive Science has extensively documented our tendency for interpolation and pattern recognition. Works like The Handbook of Imagination and Mental Simulation by Markman and Klein, or even "How Creativity Works in the Brain" by the National Endowment for the Arts all attest to this.
When artists create, they draw from their experiences, their knowledge, their understanding of the world - a process overwhelmingly of interpolation.
Now, I can see how you might be confused about my reference to ChatGPT being "super-intelligent". Perhaps "hyper-competent" would be more appropriate? It has the ability to generate text that appears intelligent because it's interpolating from a massive amount of data - far more than any human could consciously process. It's the ultimate pattern finder.
And that, my friend, is my version of "publications on the subject of how minds work." I may not be an illustrious scholar, but hey, even a clock is right twice a day! And who knows, maybe I'm on to something after all.
I have no idea how this impacts the encodability of the license from FB which may rely on things other than copyright, but as of right now, the output absolutely cannot be copyrighted.
That being said, isn't a prompt guidance?
If they are, and the license allows creating finetuned models but not using the output to improve the model, then the derived model is not a violation, but it might be a derivative work.
Your customers bought that product under license A. Afterwards it turned out that you pirated some artwork from disney. Then your customer can sue you (not disney) to make things right. The specific license of the original work seems quite irrelevant here.
edit: you seem to see the customer as the primary victim here instead of Disney, but if Disney weren't a victim the customer wouldn't have a case.
So, no matter what they TOS says, it's not an infringing work.
> Downstream users can't just integrate the derivative work into a product without abiding by the GPL terms
You absolutely could do this if the original work is not protected by copyright, or if you use it in a way that is transformative and fair use.
Similarly, if someone creates a fair use/transformative work then the license can also be ignored.
That's... not really accurate. See the concept of tortious interference with a contract.
That said, doing this inside one company or with subsiduaries probably wouldn't fly.
Can you summarize why weights would not be copyrightable or give me pointers to sources that support that view.
1. Model weights are the output of mathematical principles, in the US facts are not copyrightable, so in general math is not copyrightable.
2. Model weights are the derivative work of all copyrighted works it was trained on - in which case, it would be similar to creating a new picture which contains every other picture in the world inside of it. Who is the copyright owner? Well, everyone, since it includes so many other copyright holders' works in it.
related, there was a presentation (i've lost the reference) on automatic song (tune?) generation where the presenter claimed (rather humourusly) that he'd generated all the songs that had ever been and will ever be so that while he was infringing on a large but finite number of songs, he was non infringing on an infinite number of future songs. So, on balance he was in a favourable position.
One cannot hold copyright facts, but one can "copyright" a collection of facts like a search index or a map.
In the case of an LLM, I don't think that the work of compiling the training data probably would qualify by analogy to the phonebook example.
But on reflection, you are totally right, I was just getting mixed up on the distinction between copies and the creative works themselves. Machine output of something is generally just a copy of something. Whatever it is a copy of may be a copyrightable work, and if so, whoever came up with that original work has the right to all the copies output by machines (or copies generated by hand-tracing, or whatever).
Anyway, on LLMs... Even if we assume LLM weights are just copies (machine outputs) of whatever inputs they were trained on, then I assume I would automatically own the exclusive right to restrict the distribution of weights of a 'Me' chatbot trained exclusively on my own writings. But what if someone else comes along and writes a load of bespoke code specifically to generate improved weights for this same model, so the resultant chatbot works much better in conversation (still with my tone of voice, but with better performance and better interpretation of questions)? Is that programmer not adding some creative value, such that we might both have a right to restrict distribution of those improved weights? (NB. it's common for an item to be a 'copy' of multiple original works, e.g. copies of Jimi Hendrix's cover of Bob Dylan's 'All Along the Watchtower'.)
If it isn't recognisable, then it's merely _distributed_ plagiarism. A million output, each of which are 0.0001% plagiarising each of million inputs.
Is The War on Drugs a VC-funded band replacement?
Are other future bands going to learn from The War on Drugs?
https://www.cbsnews.com/news/ai-stable-diffusion-stability-a...
https://www.documentjournal.com/2023/05/ai-art-generators-mo...
If the original training data is a copyrightable (derivative or not) work, perhaps eligible for a compilation copyright, the model weights might be a form of lossy mechanical copy of that work, and be both subject to its copyright and an infringing unauthorized derivative if it is.
If its not, then I think even before fair use is considered the only violation would be the weights potentially infringing copyrights on original works, but I don’t think incomplete copy automatically works for them the way it would for an aggregate; I’d think you'd have to demonstrate reproduction of the creative elements protected by copyright from individual source works to make the claim that it infringed them.
A codec conversion is not copyrightable. The original song which is still present enough in the conversion to impact its ability to be distributed, is still copyrightable. But you don't get some kind of new copyright just because did a conversion.
For comparison, if you take a public domain book off of Gutenberg and convert it from an EPUB to a KEPUB, you don't suddenly own a copyright on the result. You can't prevent someone else from later converting that EPUB to a KEPUB again. Copyright protects creative decisions, not mathematical operations.
So if there is a copyright to be held on model weights, that copyright would be downstream of a creative decision -- ie, which data was it trained on and who owned the copyright of the data. However, this creates a weird problem -- if we're saying that the artifact of performing a mathematical operation on a series of inputs is still covered by the copyright of the components of that database, then it's somewhat tricky to argue that the creative decision of what to include in that database should be covered by copyright but that copyrights of the actual content in that database don't matter.
Or to put it more simply, if the database copyright status impacts models, then that's kind of a problem because most of the content of that training database is unlicensed 3rd party data that is itself copyrighted. It would absolutely be copyright infringement for OpenAI/Meta to distribute its training dataset unmodified.
AI companies are kind of trying to have their cake and eat it too. They want to say that model weights are transformed to such a degree that the original copyright of the database doesn't matter -- ie, it doesn't matter that the model was trained on copyrighted work. But they also want to claim that the database copyright does matter, that because the model was trained on a collection where the decision of what to include in that collection was covered by copyright, therefore the model weights are copyrightable.
Well, which is it? If model weights are just a transformation of a database and the original copyrights still apply, then we need to have a conversation about the amount of copyrighted material that's in that database. If the copyright status of the database doesn't matter and the resulting output is something new, then no, running code on a GPU is not enough to grant you copyright and never really has been. Copyright does not protect algorithmic output, it protects human creative decisions.
Notably, even if the copyright of the database was enough to add copyright to the final weights and even if we ignore that this would imply that the models themselves are committing copyright infringement in regards to the original data/artwork -- even in the best case scenario for AI companies, that doesn't mean the weights are fully protected because the only copyright a company can claim is based on the decision of what data they chose to include in the training set.
A phone book is covered by copyright if there are creative decisions about how that phone book was compiled. The numbers within the phone book are not. Factual information can not be copyrighted. Factual observations can not be copyrighted. So we have to ask the same question about model weights -- are individual model weights an artistic expression or are they a fact derived from a database that are used to produce an output? If they're not individually an artistic expression, well... it's not really copyright infringement to use a phone book as a data reference to build another phone book.
Its a mechanical copy subject to the copyright on the original, though.
https://en.wikipedia.org/wiki/Integrated_circuit_layout_desi...
Let’s talk about more complex models. What if my model shares 5% of the same weights with your model? What about 50%? What about 99%? How much do these have to change before you’re in the clear? What if I take your exact model and run it through some extra layers that don’t do anything, but dilute the significance of your weights?
It’s a murky area, and I’m inclined to think copyright is not at all the right tool to handle the legality of these models (especially given the glaring irony they are almost all trained using copyrighted material). Patents, perhaps better suited, but I’m also not sold.
What if there were billions of knobs, tuned after years of feedback and observations of the sound output?
If they can't charge for and control those other things, then we'll likely see far fewer companies releasing weights. Most of this stuff will move behind APIs in that scenario.
> A century later, in Feist Publications v. Rural Telephone Service Co., the Supreme Court confirmed that originality is a constitutional requirement
I don’t see why, for-profit companies release permissively-licensed ooen-source code all the time, and noncopyrightable models aren't practically much different than that.
As mercenary as it may sound, what these companies are trying to do is find a business model that is as friendly to themselves as it is hostile to their competitors.
This is all part of the jockeying.
> I don’t see why, for-profit companies release permissively-licensed ooen-source code all the time
I agree with this - however, they tend to open-source non-core components - Google won't release search engine code, Amazon wont release scalable-virtualization-in-a-box, etc.
I'm confident that Facebook won't release a hypothetical Llama 5 in a manner that enables it to be used to improve ChatGPT 8 - the aim will be unchanged from today, byt the mechanism will shift from licensing to rate-limiting, authentication & IP-bans.