Can you summarize why weights would not be copyrightable or give me pointers to sources that support that view.
Can you summarize why weights would not be copyrightable or give me pointers to sources that support that view.
Let’s talk about more complex models. What if my model shares 5% of the same weights with your model? What about 50%? What about 99%? How much do these have to change before you’re in the clear? What if I take your exact model and run it through some extra layers that don’t do anything, but dilute the significance of your weights?
It’s a murky area, and I’m inclined to think copyright is not at all the right tool to handle the legality of these models (especially given the glaring irony they are almost all trained using copyrighted material). Patents, perhaps better suited, but I’m also not sold.
1. Model weights are the output of mathematical principles, in the US facts are not copyrightable, so in general math is not copyrightable.
2. Model weights are the derivative work of all copyrighted works it was trained on - in which case, it would be similar to creating a new picture which contains every other picture in the world inside of it. Who is the copyright owner? Well, everyone, since it includes so many other copyright holders' works in it.
related, there was a presentation (i've lost the reference) on automatic song (tune?) generation where the presenter claimed (rather humourusly) that he'd generated all the songs that had ever been and will ever be so that while he was infringing on a large but finite number of songs, he was non infringing on an infinite number of future songs. So, on balance he was in a favourable position.
One cannot hold copyright facts, but one can "copyright" a collection of facts like a search index or a map.
In the case of an LLM, I don't think that the work of compiling the training data probably would qualify by analogy to the phonebook example.
But on reflection, you are totally right, I was just getting mixed up on the distinction between copies and the creative works themselves. Machine output of something is generally just a copy of something. Whatever it is a copy of may be a copyrightable work, and if so, whoever came up with that original work has the right to all the copies output by machines (or copies generated by hand-tracing, or whatever).
Anyway, on LLMs... Even if we assume LLM weights are just copies (machine outputs) of whatever inputs they were trained on, then I assume I would automatically own the exclusive right to restrict the distribution of weights of a 'Me' chatbot trained exclusively on my own writings. But what if someone else comes along and writes a load of bespoke code specifically to generate improved weights for this same model, so the resultant chatbot works much better in conversation (still with my tone of voice, but with better performance and better interpretation of questions)? Is that programmer not adding some creative value, such that we might both have a right to restrict distribution of those improved weights? (NB. it's common for an item to be a 'copy' of multiple original works, e.g. copies of Jimi Hendrix's cover of Bob Dylan's 'All Along the Watchtower'.)
If it isn't recognisable, then it's merely _distributed_ plagiarism. A million output, each of which are 0.0001% plagiarising each of million inputs.
Is The War on Drugs a VC-funded band replacement?
Are other future bands going to learn from The War on Drugs?
https://www.cbsnews.com/news/ai-stable-diffusion-stability-a...
https://www.documentjournal.com/2023/05/ai-art-generators-mo...
If the original training data is a copyrightable (derivative or not) work, perhaps eligible for a compilation copyright, the model weights might be a form of lossy mechanical copy of that work, and be both subject to its copyright and an infringing unauthorized derivative if it is.
If its not, then I think even before fair use is considered the only violation would be the weights potentially infringing copyrights on original works, but I don’t think incomplete copy automatically works for them the way it would for an aggregate; I’d think you'd have to demonstrate reproduction of the creative elements protected by copyright from individual source works to make the claim that it infringed them.
A codec conversion is not copyrightable. The original song which is still present enough in the conversion to impact its ability to be distributed, is still copyrightable. But you don't get some kind of new copyright just because did a conversion.
For comparison, if you take a public domain book off of Gutenberg and convert it from an EPUB to a KEPUB, you don't suddenly own a copyright on the result. You can't prevent someone else from later converting that EPUB to a KEPUB again. Copyright protects creative decisions, not mathematical operations.
So if there is a copyright to be held on model weights, that copyright would be downstream of a creative decision -- ie, which data was it trained on and who owned the copyright of the data. However, this creates a weird problem -- if we're saying that the artifact of performing a mathematical operation on a series of inputs is still covered by the copyright of the components of that database, then it's somewhat tricky to argue that the creative decision of what to include in that database should be covered by copyright but that copyrights of the actual content in that database don't matter.
Or to put it more simply, if the database copyright status impacts models, then that's kind of a problem because most of the content of that training database is unlicensed 3rd party data that is itself copyrighted. It would absolutely be copyright infringement for OpenAI/Meta to distribute its training dataset unmodified.
AI companies are kind of trying to have their cake and eat it too. They want to say that model weights are transformed to such a degree that the original copyright of the database doesn't matter -- ie, it doesn't matter that the model was trained on copyrighted work. But they also want to claim that the database copyright does matter, that because the model was trained on a collection where the decision of what to include in that collection was covered by copyright, therefore the model weights are copyrightable.
Well, which is it? If model weights are just a transformation of a database and the original copyrights still apply, then we need to have a conversation about the amount of copyrighted material that's in that database. If the copyright status of the database doesn't matter and the resulting output is something new, then no, running code on a GPU is not enough to grant you copyright and never really has been. Copyright does not protect algorithmic output, it protects human creative decisions.
Notably, even if the copyright of the database was enough to add copyright to the final weights and even if we ignore that this would imply that the models themselves are committing copyright infringement in regards to the original data/artwork -- even in the best case scenario for AI companies, that doesn't mean the weights are fully protected because the only copyright a company can claim is based on the decision of what data they chose to include in the training set.
A phone book is covered by copyright if there are creative decisions about how that phone book was compiled. The numbers within the phone book are not. Factual information can not be copyrighted. Factual observations can not be copyrighted. So we have to ask the same question about model weights -- are individual model weights an artistic expression or are they a fact derived from a database that are used to produce an output? If they're not individually an artistic expression, well... it's not really copyright infringement to use a phone book as a data reference to build another phone book.
Its a mechanical copy subject to the copyright on the original, though.
https://en.wikipedia.org/wiki/Integrated_circuit_layout_desi...
If they can't charge for and control those other things, then we'll likely see far fewer companies releasing weights. Most of this stuff will move behind APIs in that scenario.
What if there were billions of knobs, tuned after years of feedback and observations of the sound output?