Machines and algorithms are not legally recognized as being able to author original non-derivative works.
> put a human in the place of the LLM
But also, no, if you have a team of humans doing rote matrix multiplication instead of an LLM, that does not make it so the matrix multiplication removes copyright. Also, at this point LLMs require so much math that you can't replace them with humans, even if the humans have quite fast fingers and calculators.
"Machines and algorithms are not legally recognized as being able to author original non-derivative works"
Neither are monkeys. This doesn't mean a monkey's painting is any more or less derivative, or any more or less subject to a copyright claim. It only means that there is not a second copyright attached to the resulting work.
Let's look at a totally different analogy: compression algorithms.
If I take a digital artist's work which they publish as a png or psd file, and I use some algorithm to convert it to a jpg file, well, I definitely transformed the work in terms of bytes. It's a smaller file, I threw out a lot of data, you can't get the original back.
Yet, this does not change the copyright in any way. A computer applied a rote transformation.
An LLM is really just a very complicated compression algorithm. It takes an input of a bunch of copyrighted works, compresses them into a model, and then uses more algorithms to uncompress them into approximations of the original ("responses")
In the image analogy, an LLM response is similar to upsizing the compressed jpg back into a png (and getting a slightly different image since the process was lossy).
Is there a way that an LLM isn't, legally, a compression algorithm for a large set of copyrighted works?
An LLM outputting a summary of someone's work (1) doesn't create a new copyright work (so no profit can be derived from it's sale) but (2) would fail the test of whether it was competing with the original copyright work.
i.e. no one looking for a summary of a comedy skit is then going to consume that in preference to consuming the original skit. If you tried to argue that was the case, you'd then have to answer why a human review or wikipedia summary does not consitute an infringement.
An LLM is a lossy compression algorithm for a body of data (which may or may not consists of multiple “works” under copyright, and any works included in the data may or may not be protected by copyright), which body as a whole likely comprises a work (as compilation) which may or may not legally be derivative of some or all copyright protected works contained in the compilation, before considering Fair Use analysis.
It is not particularly a compression algorithm for the individual works if the body of data consists of individual works.
Works produced by monkeys, like works produced by computers, cannot be copywriten. They are functionally identical in this regard. That is why it's relevant.
This is how the current law works. The analogies will break down with AGI and new law will need to be created. Where LLMs fit into this process is an open question.
feeding input to a program is pretty clearly categorically different than providing source material to a human being
“original non-derivative” is noise: only humans can author works. This is especially equally true of derivative works, which must themselves being distinct works of authorship (a mechanical copy is not a derivative work, its a copy.)