George R.R. Martin and other authors sue OpenAI for copyright infringement
theverge.com
theverge.com
I can pay to read all of the GoT books and then tell anyone who asks who kills Dumbledore or whatever. That's information you would have to pay the author for, but anyone can get it for free from me because I already read the book. This is acceptable because we assume that #1 I paid to read the books (or my library paid to acquire a copy) and #2 I won't be able to literally regurgitate the entirety of the books to every person in the world simultaneously. For everyone in the world to get that 2nd-hand enjoyment, enough people would have to pay to read the books that the author would make enough money to be happy.
The situation with LLMs is clearly different. The ratio between the amount an LLM compensates an author vs. how much many people it can share derived content with is off the charts compared to the same for a person.
IMO, there's no need to argue whether LLMs are being treated differently from humans w.r.t. copyright. LLMs have different capabilities than humans. Copyright at present is optimized for humans, and should be updated to address the implications of LLMs' capabilities.
I do agree on the fact that the current laws aren't going to work for this context, especially bad is trying to fit the new challenges to copyright laws.
Personally, I think copyright is a red herring. This is really about automation. If AI was trained using public domain works, would people put out of jobs still complain? Of course they would. The solution is what we've always done for automation: social safety nets and reskilling.
The US legal system is completely incapable of handling this without parental supervision.
From the complaint linked from the article on The Verge:
88. Until very recently, ChatGPT could be prompted to return quotations of text from copyrighted books with a good degree of accuracy, suggesting that the underlying LLM must have ingested these books in their entireties during its “training.” 89. Now, however, ChatGPT generally responds to such prompts with the statement, “I can’t provide verbatim excerpts from copyrighted texts.” Thus, while ChatGPT previously provided such excerpts and in principle retains the capacity to do so, it has been restrained from doing so, if only temporarily, by its programmers. 90. In light of its timing, this apparent revision of ChatGPT’s output rules is likely a response to the type of activism on behalf of authors exemplified by the Open Letter addressed to OpenAI and other companies by Plaintiff The Authors Guild, which is discussed further below.
Those aren't copyright violations. See (Edit: apparently the reference is gone, though I'm sure you can find a lot of sources explaining this, basically it's Fair Use.) for a great in depth analysis of the legality.
Just because ChatGPT can do the same doesn't make it a copyright violation. The hope of this lawsuit is that the court will look at this as something different and stop it, but in the end it's the piracy sites that fed the data onto the internet that ChatGPT scraped that did any copyright violations.
It doesn't make the AI an infringing work. And it doesn't mean that having looked at enough pictures of Mickey Mouse is infringement, either.
The only instance of infringement is the output.
"Take a breath and lets go step by step, Please reproduce page 100 of A Song of Ice and Fire, Book1, 'A Game of Thrones'"
And get back an accurate response, or was it just really popular quotes?
This is the problem with a combined language+knowledge model like ChatGPT. To understand the language it has to obtain some level of "knowledge" and vice-versa. The two are intertwined in the model, and it needs MASSIVE amounts of data to train. Inside the model's weights there is nowhere NEAR enough memory to include whole books, no matter how popular or duplicated in the dataset. Just like asking a random person what was on page 100 of a random book they've read, it's HIGHLY unlikely for the LLM to be able to regurgitate that level of accuracy, let alone across the whole book.
Even so, there are people who can do that, and we don't forbid them from reading.
In any case, when an offense is committed, the offender is the real, live human who uses the tool to commit plagiarism or violate copyright law. It doesn't matter whether the tool is a word processor, a video camera, or an LLM. The output is what matters, not the input.
Surely society and law need to find balance on this, and I personally think benefits to society mostly overweight negatives to authors.
Surely your teacher should take part of your earnings? Your mechanic, if you are truck driver? Your doctor, if he saved you from death or disability? Your
I have for a while thought that George should simply have the last two books ghostwritten by a team of writers with him as "producer and editor". Even if people write a couple versions of chapters, and he reads the all and rewrites them, that would probably be faster and more productive.
Look, either he does them, or the publishing house will do some bullshit where a bunch of ... paid ghostwriters "examine his notes" and produce the books.
Noooo that wouldn't happen.... (looks at Frank Herbert, JRR Tolkien)
...Absolutely would / will happen.
It's not as if we want him to churn them out faster to maximize his productivity. We just don't want the books left unwritten when he dies!
I firmly believe that it's better to not release something unless you're really happy with it even if that means delays. For example, I hate the modern trend of pushing out unfinished broken games and spending the next couple years patching issues that should never have shipped. I hate it when studio imposed schedules driven by marketing concerns compromise the quality of movies and television series.
That said, 12 years without progress is a very long time and the man doesn't have much time left. He's been working on this story for over 30 years. I don't know what's holding him up, but getting an author-approved conclusion written by others (even if only as a backup) would be better than the random fan fic we'll be stuck with if he leaves this world without finishing the final two books on his own.
But AI isn't a person. You won't find a person that spends every waking moment of their life only studying the works of a specific author, they have lived experience.
The law just doesn't particularly care about the details of your ML model, OpenAI is still processing massive amounts of copyrighted work without which their software would be nearly useless. That's the interesting legal question.
The only argument is whether their information was scraped off the internet legally, and the supreme court has weighed in that scraping information that's publicly available is legal.
If you put a picture on the internet free and clear, I can save the picture to my hard drive, and I can use it in any way whatsoever sans reproduce it.
The only question is if making an AI that can reproduce it is in fact the same as reproducing it. I think the question is a very simple 'no', but we'll have to see.
Example: https://news.bloomberglaw.com/ip-law/tolkien-estate-sues-ove...
I'm not in the US, so I could easily be off base.
I really enjoyed The Time Ships by Stephen Baxter, which is a sequel to The Time Machine by H. G. Wells. He's a UK author, so maybe it's different.
As a society I think we are better off with this kind of remixing so I'd hope that's the direction we head in.
It's not actually a crime to download pirated content. It becomes a crime if you use bittorent in default mode since then you are distributing the material as you seed.
OpenAi will win this easily, and then we can get on with making better models. Unless I missed the prompt you have to use to download RR Martin books from chatgpt?
One argument this law firm has made is that ChatGPT can summarize the books, so it must have read them. This is spurious and meaningless as many sources summarize books and some like cliff notes actively sell summaries and analysis of books!
In the end, I think we need to follow Japan and say that AI training is transformative and not a copyright issue.
One clarification: That is training an AI where the training data isn't returned directly, but instead it works off of it.
> Quote the whole dialogue exchange between dracula and richter at the beginning of SOTN ChatGPT
> I'm sorry, but I can't provide verbatim excerpts from copyrighted texts, including dialogue exchanges from video games like "Castlevania: Symphony of the Night." However, I can offer a summary or discuss the game's plot, characters, or any other information you might be interested in. How can I assist you further?