Authors sue OpenAI for using their works without proper licensing
nytimes.com
nytimes.com
This kind of luddism sees copyright as a way to enrich rights holders, as opposed to "promoting the progress of science and the useful arts".
I was onboard initially thinking we were talking about OpenAI ingesting Game Of Thrones as training material, but it appears George et al are just mad because it can make stories with their characters.
This is far from the authorship/copyright problem of AI.
Terrible headline. They're not suing for theft, they're suing for copyright infringement.
The rest of the article is reasonable, and they link to the complaint, which is something every article about a lawsuit should do.
https://fingfx.thomsonreuters.com/gfx/legaldocs/xmvjlbqbnvr/...
EDIT: i just asked an ai and it said that single quotes are a british-ism?
Either way, it will be interesting to see how this goes. Lots of weird arguments on both sides, so it will be interesting to see what rulings we get.
https://ia904703.us.archive.org/10/items/gov.uscourts.nysd.6...
Named plaintiffs include David Baldacci, John Grisham and Scott Turow.
Seems like this lawsuit could set precedence that the future I describe would not be allowed.
It's legal for me to make digital copies of copyrighted works I own so long as they're not redistributed. If I have a local AI maintaining my personal data for reasons other than generating competing works, that shouldn't involve copyright at all.
What specifically hurts OpenAI is the monetization and commercial benefit from other people’s work. They also fall clearly in the space of civil claims of copyright violation, if not criminal (however the use of models to produce copyright materials is likely criminal infringement). (IANAL, but a law professor friend made these claims to me, YMMV)
And it'll be cross border, so even more difficult to enforce.
A LLM does the same thing but instead of search results it emits a stream of tokens.
Win/win?
Ideas have never been the scope of copyright and it wasn’t in its democratic mandate. If creatives want that change, fine, advocate for a change of the law
This isn't about ideas, it's about a specific individuals work given that the reproduced text lifts literal characters out of Martin's book. That has always been covered by IP law. Canonical example, you cannot write a novel about Harry Potter, you can write a book about a wizard going to a magical school.
I don't think so. It's not illegal to look at or learn from copyrighted materials. If you start producing the materials it becomes a different question. I think the same applies to AI.
Or disagreement anyway, about how comparable photocopiers & copyright are to generative models and protection from unauthorized automated style reproduction.
How I look at it:
1. In both cases, reproduced copies or reproduced styles, automation destroys economic incentives for creators to make any sustained effort.
Without economic protection, it isn’t even a question of less motivation. Creator’s like everyone else need to eat.
2. So we protect creative works from complete copies in order to have more creative works.
And it is primarily about automation and mass reproduction.
Nobody is worried about people hand copying Atlas Shrugged.
3. But we also protect copyrighted works from partial copying.
Only copying chapters 1-3? Not allowed
Only copying the plot but changing all the names, locations, fashion and colors? Not allowed.
4. So now it turns out a different substantial part of a work can be copied via automation. It’s style.
Well if you can protect a works plot from automated copies, why not a works style?
It is a substantial piece of a creative work.
Reasons for protecting style come down to protecting any major part of a copyrighted work.
The only thing different now is we have “style reproducers”.
So we have to decide, is this essentially the same situation as copyright addresses, or not?
5. It is.
The exact same trade offs between protection and incentivization exist for extracted & mass reproduced style as they do for extracted and reproduced plot.
Was Sword of Shannara pretty derivative of Tolkien? Yeah. But I assume it was pretty far from a copyright violation.
Movies are sued all the time for copyright infringement due to substantially copying plot and character elements. [0]
Because these cases tend to each be unique, the line between infringement and non-infringement gets settled very much on a case by case basis.
As a result of this inherent unpredictability, most cases involve the accused settling with the aggrieved party to get the lawsuit dismissed.
This is common in many areas of civil law.
A few examples:
1. *"The Island" (2005)* - Accusation: Similarities to the 1979 film "Parts: The Clonus Horror." - Outcome: Settled out of court. [1]
2. *"Frozen" (2013)* - Accusation: Claimed similarities to a short film named "The Snowman." - Outcome: Disney settled the case. [2]
3. *"Coming to America" (1988)* - Accusation: Art Buchwald claimed the movie was based on his script. - Outcome: Paramount settled for an undisclosed amount. [3]
4. *"The Terminator" (1984)* - Accusation: Harlan Ellison claimed it was similar to an episode of "The Outer Limits." - Outcome: Settled out of court, and an acknowledgment was added to later copies. [4]
5. *"Disturbia" (2007)* - Accusation: Accused of being similar to Alfred Hitchcock's "Rear Window." - Outcome: Initially dismissed, but a settlement was reached. [5]
[0] https://movieweb.com/movies-accused-of-copyright-infringemen...
[1] https://en.m.wikipedia.org/wiki/Parts:_The_Clonus_Horror
[2] https://ew.com/article/2015/06/25/disney-frozen-lawsuit-the-...
[3] https://en.m.wikipedia.org/wiki/Buchwald_v._Paramount
[4] https://www.cbr.com/terminator-harlan-ellison-credit/
[5] https://www.flixist.com/new-disturbia-and-rear-window-lawsui...
Stranger things takes liberally from a lot of Stephen King as Spielberg elements not outright but in spirit and tone, why isn't Stephen King suing the Duffer brothers for reading his shit and coming up with ideas for books based on that?
See my reply to a sibling comment of yours.
Personally I currently feel that (at life +70 years) the copyright pendulum has gone too far towards the rights of publishers (not necessarily authors) as is.
That said, I'm open to good arguments to change my mind. Why do you feel that authors should be given this additional right to control what is used for AI training? What would be the public good or public trade-off here?
What copyright buys is that no one else can distribute verbatim copies of large amounts of your work. But a lot of other uses are allowed.
It covers the expression of ideas. Which in the case of a book is mostly the text as written. And, yes, doing some substitution of character names etc. may still violate copyright but you certainly can’t keep me from writing an article about the main points you make in your book.
Consider the aesthetic landscape where creators do not have control over whether their work is used to train an AI versus one where they do. It's hard to predict with certainty, but my model is this:
No control: Anyone's work is fair game to be trained. If I want to make a prompt of "A graphic novel in the visual style of Moebius, written by Stephen King, set in Westeros" I can get something based on King and Martin's actual words and Moebius' actual drawings, without compensating them. Neat! However, potential new novelists see that quality novels can just be churned out for free or low cost and so, actually sitting down to write a new novel becomes a niche, geek thing to do. There's no money in it. These new novels just get thrown into the ml bin, fodder for the next version.
With control: novelists and other creators know they can make money from their work because they can make business decisions about how and when their work trains a model. We all get to see more new, professional quality creativity. Those who want to read Conan as written by Lord Dunsany can still see that, since those works are in the public domain.
In a sense it's like driving down the highway with a duffel bag of cannabis flower in a state where possessing and traveling with few ounces is no problem--something commercial is probably happening. Why is that prohibited? Perhaps another debate, but just trying to connect the implied intent aspect.
If an AI were being trained for strictly academic reasons then I'd agree with fair use and that type of arguments. But if the AI itself has a subscription fee, then whoever is subscribing is also probably using the work for real or anticipated commercial gain. Hence investing money.
True that hobbyists spend money with no intention to gain commercially, and we may do that at a higher than average rate as tech workers because we have usually have a decent amount of excess money from our work. But money is pretty scarce to most people and businesses with set non-investment budgets, so if they're spending it on AI there's little doubt it's with commercial intent.
So in conclusion, I do think there's both merit to the authors' case related to intent to commercialize and room for doing unlicensed non-commercial AI training.
This is a good thing. We're going to see an explosion of indie games, movies, and more that never could've been made before by a single, dedicated person.
If not, how are we going to legally codify the difference between that and an LLM?
Here are screenshots of the completion and the diff with the real book.
When I ask it for anything from Game of Thrones it refuses beyond offering a summary.
But this is all begging the question.
How are you going to codify that? "You can't feed my text into an algorithm that has the ability to reproduce it"?
I think for a few high-profile authors (for books like Harry Potter and Game of Thrones), OpenAI probably installed some output filters in order to not get sued too hard. Of course, I can't definitively check that without access to the raw model. Which OpenAI conveniently doesn't provide.
Thought experiment:
What if every person who dowloads materials from the above sources claimed that they were doing so only to "train AI".
Many such persons who download from those sources are probably doing so for noncommercial purposes, for example, academic research. Whereas, according to this compaint, OpenAI "intend[s] to earn billions from this technology."
More discussion days back when this was news:
So AI works can't be copyrighted but training AI using copyrighted materials are copyright infringement?
The human owner/operator did not have proper licensing to perform this action, trying to argue "but with an AI" doesn't change the act.
The law is weird.
I'm wondering what will the authors do if we develop AIs that are able to find new artistic styles that are not in the dataset ? Would it still pose problem to use their content to learn how NOT to imitate them ?
Seems it is possible in collaborative filtering:
> Yes, in collaborative filtering, finding empty classes is possible. To recommend items for these gaps, utilize adjacent class information or employ techniques like matrix factorization, content-based filtering, or hybrid systems. These methods predict preferences based on observed patterns, similarities between items, and user preferences, filling in missing data.