Major U.S. newspapers sue OpenAI, Microsoft for copyright infringement
axios.com
axios.com
I would see this as strengthening, rather than undermining, the argument. The AI companies in those agreements are paying for something, which implies that the work both has value, and should be paid for.
I think it’s time to start managing our expectations around GPT-5
Every author ever has learned from, been inspired by and influenced by the books they've read before. Same for every poet, painter, musician, sculptor. So much so that it's a common interview question to ask such people about their influences.
But somehow, if it's AI instead of a human, the same behavior is supposed to be copyright infringement in the eyes of the greedy copyright cartel? Preposterous.
You can't go to theater and instruct your computer(a.k.a. camera) to consume & learn or summarize the movie/performance on your behalf. If that's possible and allowed, I'm sure the admission fee will be drastically different.
Google will not let you instruct your computer to consume all the search results for a keyword and filter out good results & summarize them (you can technically still do that but Google will try their best to make sure you don't). If you do, they'll throw you into a hell that is infinite captcha loop. Yes, they are hypocritical to do that but that is a different issue.
A human and a computer are not comparable either in capabilities or their lifetimes. Let's not mix them up.
Fair Use is the main doctrine relevant in this situation, and does not inherently distinguish between manual and automated creation. Use of traditional algorithms, like the caching/thumbnailing/snippeting done by search engines, is already covered by it.
There are cases where "but it's not human" makes sense because the relevant law does make that distinction (e.g: registering a work for copyright in the US depends on human authorship), but I think it's a fairly weak point in this case. You could instead argue that it's not solely "doing something a human can do" and how that impacts on the Fair Use factors.
LLMs are not AI. LLMs have nothing to do with how the human brain works, they have at most, a passing resemblance to how a brain could work, with the tech we have available. They aren’t human. They are code, and the humans who create these programs cannot hide behind “yerm acktchually it’s a brain so your laws don’t apply”
On that note, I would like someone to point out which bits in the model file contain the copy that's infringing.
Copyright is not some kind of divine law handed down by god.
It is exactly what copyright is for.
Want proof that it’s useful?
> “it would be impossible to train today’s leading AI models without using copyrighted materials“ - OpenAI
Being unable to survive if you had to enter into mutually consensual agreements with your suppliers is a pretty good sign that you’re rent-seeking or stealing.
1) Infringement is not theft and is usually handled differently. The (poorly-named imho) NET act, which criminalized some non-commercial infringement, had to do so explicitly.
2) Sometimes exclusive/monopoly rights are not in the public interest, and compulsory licenses are desirable.
2) Supports "rent-seeking" by providing an example of legislative counterbalance.
>Being unable to survive if you had to enter into mutually consensual agreements with your suppliers is a pretty good sign that you’re rent-seeking or stealing.
Did you get permission from OpenAI to share that quote or are you rent-seeking/stealing?
I think you'd agree that there are more reasons than just that to not always enter into agreements.
Contrast with the entire value prop of LLMs being that you can utilize the knowledge in them without paying any time or credit toward the source, which does indeed destroy all commercial and most non-commercial incentives of producing and sharing such information.
The entire point of copyright law is to protect the incentives to produce work. The models are certainly doing something of value that also should be protected (and surely they’ll utilize law to do so), but ultimately a system that sucks up all prior works and obviates the need to view/buy them and destroys almost all incentive to produce new ones will not stand.
Of course this will boil down to model creators saying it doesn’t create incentive (except for “possibly capturing all future value of the light cone,” when talking amongst themselves) and original creators claiming that it does.