I'm not sure OpenAI offering a service for profit falls into that category.
I'm not sure OpenAI offering a service for profit falls into that category.
In the case of LLM training, it's for the same purpose as the source material -- to generate code, or writing, or photographs, etc. Not only that, but in several instances it's been shown to reproduce source material, which is either derivative work or straight copying, depending.
They're different situations.
If it's used in a reference/"inspiration" capacity (as opposed to verbatim copying), I doubt the rightsholder have anything to stand on here. Sure, their works might have been used to make other competing works, but all art is derivative, and I don't see why it would be legal for a human artist to "train" on past works of art but not AI.
Alleging that AI models can reproduce some works verbatim is probably the stronger argument, but AFAIK you have to coax them pretty hard to do so, and therefore AI companies might be able to argue they're tools like photocopiers or such. Likewise, you can probably extract an entire book off google books by bruteforcing common ngrams to get the entire book, but google wouldn't be held liable for that.
The "all art is derivative" line is essentially something people try to convince others of to justify breaking copyright law. It's not grounded in reality or law. It devalues creative work by implying the machine, with no lived experiences, is doing the same thing. And it's also completely wrong about what the specific term derivative work actually means in the context of copyright.
Derivative works deal in specific. If your LLM reproduces a substantial portion of the story beats from Jurassic Park, you can bet it'd wind up in court. If it reuses identifiable characters, that is usually gonna be derivative unless it can otherwise qualify under an exemption.
"But fanfic, fanart, etc." Is a common counterpoint but misses the commerce aspect of it. Here Open AI and similar are offering paid services based upon harvesting all of this information. When they produce for you the response to the prompt, they are, effectively, distributing that to you for money. That's the point at which it becomes a problem.
As an aside, it's an act of drinking the LLM Kool-Aid to believe it can be "inspired".
> Alleging that AI models can reproduce some works verbatim is probably the stronger argument, but AFAIK you have to coax them pretty hard to do so, and therefore AI companies might be able to argue they're tools like photocopiers or such.
They can try that argument but it'll fall flat when you consider that a photocopier is reproduction agnostic, while LLMs generally have a ton of work going into them to prevent them from outputting damaging things (and they still fail). That fact makes them not at all comparable to a photocopier, setting aside the more obvious "subscription software service" different.
Also, you "know" pretty wrong about the effort required. For a recent example, see: https://www.latimes.com/entertainment-arts/business/story/20...
Here a number of people noticed getting specific producer tags basically unaltered in the output when just asking for songs of a certain genre, which then also often sound similar to existing songs.
"I want a black and white logo in the style of an 1960's Archie comic for an ice cream shop named 'Bettys'"
Once I say "1960's Archie comic", why doesn't the work instantly become derivative whether a human does it versus a computer?
If I understand your argument correctly, the person from Fiverr will not pay license fees to the owner of Archie Comics, even though he may use it as reference material.
I mean, if you ignore all the massive differences in paying a human to do something versus paying an LLM service to do it, sure. But you're effectively throwing at least ethics and care for the environment out the window in one case.
> Once I say "1960's Archie comic", why doesn't the work instantly become derivative whether a human does it versus a computer?
It does. Just because you can commission art from someone doesn't mean you won't get sued if you start trying to use it as the logo for your business. If you put Foghorn Leghorn on the logo of your chicken business, you'll be sued. Having an artist simply make you a logo like that on commission, if not transformative, could get them sued, though by doing it on commission the terms likely mean the requestor is the one who's liable
Earlier I noted clean room implementations. The software industry went to incredible lengths to be able to interoperate with competitors without violating copyright.
I agree with that. But then at what point does AI output becomes transformative versus not? No one owns the 1960's comic art style. But did mentioning "Archie" somehow make a difference? I don't think it does. I might be wrong.
So I don't understand how it becomes a problem once a computer does it versus a human. If it's a legal issue, then I might be more persuaded. If it's an ethical issue only, well... this can be thrown on top of the heap of ethical issues businesses have long ignored -- and your arguments is screaming into the wind, as it were.
I think if you argue for UBI, or some kind of remuneration then we can argue about who deserves it. Truck Drivers who lose there jobs to AI? Open Source Software Engineers? Artists? Writers? Normally this would normally be served by things like unemployment insurance in the US, but we as a society hate freeloaders.
I think with UBI or something similar we could have more people doing things they enjoy, so maybe we would have more artists, rather than less. But again, that would be, socialism, which we also hate.
And what about my other point about coaxing google books to give you a full copy of a book via multiple snippets?
>Here a number of people noticed getting specific producer tags basically unaltered in the output when just asking for songs of a certain genre, which then also often sound similar to existing songs.
Can you provide an alternate source for this? I skimmed your link and it does not substantiate that claim.
Sure, given enough time and effort maybe a person can. That's not really relevant to the lawsuit or its details though. Is your argument here they won the lawsuit which cleared the way for mass copying and redistribution?
> Can you provide an alternate source for this? I skimmed your link and it does not substantiate that claim.
If you want to hear it yourself: https://youtu.be/_wuKZR0Pv-Q
Neither the authors nor publishers received any compensation for having their work ingested. It isn't like OpenAI went to Amazon and bought one copy of every book - they downloaded a torrent.
A key part of Google's defense was that not only was it not using the entire books to reproduce the entire book, but also that it was taking measures to prevent people from abusing Google's systems to reproduce an entire book. It's a lot of work to emphasis that the impact on the market (in other words, the fourth factor) is as minimal as practicable--and that's the crux of the analysis.
When you're instead scanning someone's stock image database to build a tool to generate stock images... the fourth factor is jumping up and down screaming at you "YOU LOSE" and your best defense is that it's not the training, it's the tool built on the training data that is infringing the copyright.
OpenAI are building a product to offer to the public for profit.
If I employ 10,000 humans to read books and provide summaries or texts "inspired by" those books, I need to pay for the copies of the books those humans read.
So was google books.
>If I employ 10,000 humans to read books and provide summaries or texts "inspired by" those books, I need to pay for the copies of the books those humans read.
IANAL, but that would be perfectly legal. Summaries aren't copyrightable, and if you can acquire the book free but legally (eg. library, borrowing from a friend, buying it from a store and then returning it), there's nothing the publisher can do.
Okay, so what happens in the interim situation? If the legislature hasn't spoken yet, is it assumed to be legal or assumed to be illegal? Or is this assumption tested on a case-to-case basis, with both sides making arguments as to why it should be treated to be legal/illegal in this specific scenario?
It's difficult to argue that, for instance, training a model on all of Frank Miller's work then prompting it to generate comic art in Frank Miller's style then selling that is fair use.
And in my opinion (which is unfortunately more controversial than I think it should be) what LLMs do is far more akin to tracing than inspired creative expression. And I think intent is relevant here. Someone using an LLM to create a product in Frank Miller's style, trained on Frank Miller's work, isn't merely trying to create something inspired by his style, so much as create a Frank Miller product without having to pay Frank Miller.
"trouble" of what nature? You'd probably face more social consequences than legal.
Even if you want to argue that fair use didn't apply to Google, it clearly applies far less to what AI is used for.
It's _copy_right. If reproducing verbatim snippets was "transformative" enough to fall under fair use, I don't see why producing whole new books would not count as "transformative" enough. Copyright is a regime to grant monopoly over a specific work, it's not a regime to prevent competition from others in general.
If Google was selling brand new books created only by taking snippets from other books, that would also fall under fair use?
Google benefited from the exact same kind of bulk copyrighted data collection. They made verbatim copies of the text of both web sites and just about every book in existence!
This kind of argument seems disingenuous to me. Either ban Internet search or acknowledge that training an AI on copyrighted text is no different than a student reading every book in a public library.
Speaking of which: We all have free access to GPT 4o without advertising. It feels like asking a knowledgeable librarian.
For comparison, I had to fight for a year to get copyright permission to show book cover artwork in a library enquiry system! If I simply Google the same book titles or ISBNs, Google will show me the pictures directly. E.g.: https://www.google.com/search?q=greg+egan+eon&udm=2
How is that legal!? We had to pay to get access! In public and school libraries!
The law in most western countries is very clear that book covers are "entire" works of art, and can only be displayed by organisations that pay the copyright holders.
Google, Bing, and others violate copyright on a mass scale on a daily basis. Not to mention YouTube, TikTok, and Reels, all of which are packed wall-to-wall with "movie clips" and "TV show highlights". They're publishing copyrighted content uploaded by random people and then distributing the advertising revenue to the copyright violators instead of the copyright holders.
This isn't "caching" or "indexing", it's verbatim serving.