While harder to do as a human, if memorised a copyrighted book and then did a live reading on TV, or produced replicas from memory and sold them (the most comparable example), I’d be sued.
Humans produce derivative work all the time, and it’s fine for LLM’s to do that, but you can’t do it verbatim.
This is not the most comparable example, because it's not what ChatGPT is doing. The most comparable example is if you were hired as a contractor and the employer asked you to write verbatim some copyright content you'd memorised. If the employer then published it, they'd be the one liable, not you.
>Humans produce derivative work all the time, and it’s fine for LLM’s to do that, but you can’t do it verbatim.
Nobody's suggesting preventing humans from consuming any copyrighted content just because in future they might recite some of it verbatim, but that's what NYT want for LLMs.
No, you'd both be liable. You are not allowed to create copies of a copyrighted work, even from memory, for any commercial purpose. Making it public or not is irrelevant.
This is more obvious with spftware: if I copy a version of AutoCAD that my previous employer bought and sell it to another company, or even just use it for my current employer without showing it to anyone else, I am violating the copyright on that software, and I am liable. Even though obviously no "publishing" happened.
Similarly, if you hire a decorator to paint Mickey Mouse on the inside walls of your private kindergarten, the decorator is violating Disney's copyright just as much as you are, even if neither of you has made that public.
That's the point at which infringement occurs in your example. It's not the memorizing that's the infringement, it's the reproduction from your memory.
We shouldn't be regulating your hippocampus encoding the book, but your reproducing the book from that encoding.
Similarly, we shouldn't be regulating the encoding of material into the NN, but the NN spitting back out the material.
You'd get the same problem with someone with a photographic memory who a group of people would turn to recite them the news instead of buying the newspaper.
As of now public performance of copyrighted material is infringement.
I fully agree with the perspective that infringement in usage needs to be limited even if I strongly disagree that training is infringement.
Are they all owned by one mega-corporation, which is going to do as capitalism does, and use them to squeeze money out of all of us? Then I'm happy to ban them.
The opportunity cost of holding this technology back is going to literally be millions of people's lives given current trends in its emerging applications.
Police usage, not training.
Why should the law treat a LLM in a body reading NYT on a tablet differently than a LLM browsing the content from a website online and reading that?