I’d argue that that’s not really how the law is supposed to work.
Big media companies don't want that, zillions of small creators don't want that, the public at large has no great affection for giant tech companies getting even more power.
Big Media isn't worried about you drawing Shrek or Donald Duck at home by hand or by GPT. They won't like you distributing it profitably, but they already have tools against that. They'll also like few billions coming their way but that everyone likes.
Small Creator are worried about getting the next gig as AI may take away low end of the jobs, and they are not worried about copyright implications of the LLMs.
General Public doesn't like Tech companies but loves the products they make, nobody stopped using early Youtube or Telegram because there was piracy happening.
It will have 2 consequences:
1. No reliable law.
2. No win for humans. Only big corps (openai and corporations behind it) wins.
Is it really what we all deserve?
A: Big media companies will make billions from licensing, AI will only be produced by a few companies that can afford the licenses and creators will get an annual $10 check (see Spotify).
This really isn't clear because cognition is treated as a special exception to copyright. Every thought we have is derivative of everything we've seen before to some degree; reading a book makes our brains a derivative work. But we recognize that cognition is special.
With machines we tend to apply a strict test: Did copyright go in? If so, the output is almost certainly derivative.
With human brains, with cognition, it isn't enough to prove that a person has consumed a copywitten work prior to having a thought -- instead we judge every thought individually as to its originality.
If we are in a position to apply similar cognitive rules to an LLM then the weights won't be derivative works and we will judge each output as to its originality rather than simply assume.
Actually, no. It's considered a transformative use. If you memorize a copyrighted play or piece of music and then perform in in public, that's a copyright violation. It's the literalness of the copy that matters.
The new play is judged as to its originality.
People who have seen a play (everybody) are allowed to write new plays which aren't beholden to the copyright of the first play they've ever watched.
> This really isn't clear because cognition is treated as a special exception to copyright.
Human cognition; not the latest algorithms and their output, which some enthusiastic software engineers eagerly confuse for cognition. It's actually pretty clear.
The open question is how to handle machines that mimic the process.
It's not really an open question, except for software engineers who've talked themselves into thinking of humans as computers. A machine is not a human mind, so does not benefit from the legal exceptions and rights granted to the latter.
"A machine is not a human mind, so does not benefit from the legal exceptions and rights granted to the latter."
Five years ago this was true. It likely will not be true, eventually. The only question is where we are in this process.
Huh? You brought up "cognition," but now that has nothing to do with it?
> It's a legal question regarding large scale compilation of data.
The legal question of does "copyright goes away if your violation is big enough?"
>> "A machine is not a human mind, so does not benefit from the legal exceptions and rights granted to the latter."
> Five years ago this was true. It likely will not be true, eventually. The only question is where we are in this process.
Because of the meme magic of enthusiastic software engineers will end the "carbon chauvinism" of humanity forever?
Yeah, right.
This is just a legal fact. It has nothing to do with how an LLM operates internally, or whether an LLM is at all similar to a human mind in terms of internal mechanics.
> "The legal question of does "copyright goes away if your violation is big enough?"
Don't be fatuous.
> "Because of the meme magic"
No, because of the way the law works.
2) even if they were somehow proven to be the same there is still no reason why the same standards need to be applied to computer programs and humans because computer programs do not have any rights or legal protections.
3) cognition is not a "special exception to copyright" because it is entirely unrelated. "Copy" "right" is who has rights to make copies. Your thoughts are not considered copies because they are intangible.
4) we do not "judge every thought individually as to it's originality" because other peoples' thoughts are entirely opaque. Nobody is judging your thoughts, and if you think they are you need to take your medications.
This is false. The LLM's entire purpose is to mimic cognition.
You could argue that the operation differs in important ways - of course. But the similarity of output is literally the entire point.
"2) even if they were somehow proven to be the same"
I didn't suggest they need to be the same, proven or otherwise. I think you're not understanding. The point is that the function is similar.
How it works doesn't necessarily matter.
"3) cognition is not a "special exception to copyright" because it is entirely unrelated. "
False as a matter of law.
"4) we do not "judge every thought individually as to it's originality" because other peoples' thoughts are entirely opaque."
Also false as a matter of law. When you publish your thoughts - your works, writing, whatever they are judged as to their originality if the question of who owns the copyright is raised.
"Nobody is judging your thoughts, and if you think they are you need to take your medications."
There's no need to be snarky and disingenuous.
From the comment guidelines: Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
Purpose and mechanism are not the same thing. "Similarity of output" does not make it equivalent.
>I didn't suggest they need to be the same, proven or otherwise. I think you're not understanding. The point is that the function is similar.
Sure, go ahead and ignore all but half a sentence and then accuse me of missing the point.
>False as a matter of law.
Show me the court case where somebody was found to have violated copyright law by thinking about something.
>When you publish your thoughts
You don't publish your thoughts. You publish essays, internet comments, articles, videos, etc based on what you are thinking and those are subject to copyright law.
>There's no need to be snarky and disingenuous.
How dare you, i would never disingenuously tell somebody who thinks his thoughts belong to other people to take their psychiatric medications. Of course i did mean that they should be prescribed by a licensed physician and looking back i regret not stating that explicitly.
No one said they were. You may want to revisit my original observation.
Aside from that question, I tend to agree with the judge - LLM outputs are obviously not derivative works or copies in any sense that we normally use the phrase.
It seems important for the answer to remain no, otherwise the only entities that can afford to train on sufficient numbers of books will be big companies. No one can afford 190,000 books except huge corporations, so we’ll be surrendering our ability to train our own useful models if the price of training is that high.
190k books isn’t even that much. OpenAI probably trained on millions.
I can see the argument it might reduce open models, but how is asking for permission (and maybe sometimes paying) of the rights holders going to reduce competition? It would be a level playing field. If you're a company that's working on a LLM, you have to ask permission to ingest data that doesn't belong to you. All companies working on an LLM would have to follow these rules. So how is it anti-competitive?
I'm not sure how to interpret this line. By "violation of copyright" are you referring to the ideal copyright law that's in your mind but doesn't exist yet, or are you referring to the copyright statutes and case law that are currently in effect? And in your opinion is OpenAI violating the latter, the former, or both?