We have entire professions (musicians, actors) who develop this ability.
Is that, in your view, 'illegitimate' too?
Why or why not?
We have entire professions (musicians, actors) who develop this ability.
Is that, in your view, 'illegitimate' too?
Why or why not?
However, copyright law seems to make a distinction between expression/implementation of an idea (which can be copyrighted), and the idea itself (which generally speaking, cannot be). This is why clean room implementation is a legally sanctioned protocol to copy functionality without running afoul of copyright law.
Whether the language model is copying the "expression" (or a particular implementation) of a computating function, or whether it is just somehow re-implementing it given user's instructions, is going to be an interesting debate, and I really doubt people who think it's one way or the other has actually understood the nuances.
I'm sure you've heard of covers? Well every cover that is published affords a royalty to (at least) the original authors of both the lyrics and composition. The artist may get some money, depending on how the work was licensed.
Do you see a difference between training and reproduction?
Yes someone may teach another person based on their memory, but even if that person still performs that work they still are legally required to license the work.
Yes, if the only thing they were doing was training, sure. But it’s not. They’re training and then presenting the data and given the way LLMs are trained, there is no guarantee a transformation even takes place.
At the end of the day LLMs should be licensed under the current copyright system. Maybe OpenAI need to donate some money to a few politicians for that to change.
LLMs/Copilot/etc. copying other people's code is not legitimate for precisely the same reason. But big and powerful beats small and week.
To translate that to LLMs: training an LLM on songs (or code) is fine. It's the output and what you do with that output that might be problematic.
Only if you assume that LLMs are treated like people. Does an LLM have a right for decent working conditions? Is it illegal to make an LLM work 24h a day?
LLMs are not humans, therefore it seems reasonable to not blindly translate everything that applies to humans.
One problem I see with LLMs is that it is killing the job of many artists by using the copyrighted work of those very artists. The more the artists work, e.g. by doing truly original content that an LLM cannot generate because it was not trained on that, the better the LLMs get at killing the job of the artists.
It is a very weird situation: the LLMs need the artists in order to finish killing the artists. When that is finally done, it's not clear at all that LLMs will be able to generate new truly original content or whether they will be stuck in the world that the artists created before dying.
It's very different from, say, a calculator that replaced the job that real person had. Because the calculator can do at least as well as the human was doing, and the calculator was not feeding from the output of those humans.
Are you old enough to remember when humans planned in advance which video you would watch next?
Well if it prevents humans from producing art, of course this will happen. Doesn't mean it will be better art. I would argue that it will be worse by definition.
Very much so. But now that time is back.
Most people I know do exactly that. The time when everyone consumed TV passively is over.
https://www.youtube.com/watch?v=MAFUdIZnI5o
It doesn't really matter what my or anyone's view on it is, unless they're either a law maker or a judge.
People do not get sued for having abilities. They get sued for publishing copyrighted works. What tools they use to create or publish those works is not relevant.
The only interesting legal question in my opinion is wether the large language model itself can be considered a carrier of copyrighted information, like for example an encrypted DVD would be.
If you can come up with a prompt that itself does not contain substantial copyrighted information, and that prompt leads the model to come up with a substantial amount of copyrighted information, then in my opinion it's no different from an encrypted DVD, and as such would illegal to distribute without licensing. Also I believe it would be illegal to provide access to it over the web like OpenAI does.
IANAL, but I wouldn't hold stock in any company that publishes (access to) models that are easily goaded into producing substantial copyrighted information.
In fact, I once asked an actual lawyer about the information-theoretic questions raised by illegal/infringing content known to be on an encrypted storage device.
Her answer was "Yeahhhh, no, the Judge is probably not going to see it that way." :-)
Also, comparing a human to a service that runs on billions of dollars of infrastructure is silly.
I have AI LMs running locally on my GPU. They cost less than a photocopier or CD burner to operate. Pennies.
https://www.npr.org/2015/09/23/442761620/after-copyright-rul...
https://variety.com/lists/song-copyright-infringement-cases-...