OpenAI now tries to hide that ChatGPT was trained on copyrighted books
businessinsider.com
businessinsider.com
They're trying to avoid reproducing copyrighted text, which is a totally separate (and arguably more clear-cut) legal question. Input vs output.
This is what I've been wondering. Does Fair Use apply here at all? Sure, the models were trained on copyrighted material. But wouldn't the generative part of the AI count as transformative?
I know not to do that, though.
Seems only fair that the same should apply to GPT.
Honestly, this is starting to feel like copyright holders using the Big, New, Scary AI as a strawman to attack Fair Use.
The concept of training a model with the explicit intent of selling the output of that model is inherently different.
Not saying that it should be illegal. But it is clearly in violation of the spirit of existing copyright law, in my opinion.
They set out with the intent to make money, using copyrighted input. Seems pretty simple.
See dragonwriter's comment for the articulate version of what I'm saying
For example, clearance of music samples.
Authors do not need licenses for works that inspired them.
In the "monkey selfie" case, a photographer named David Slater set up a camera in the Indonesian jungle, and a macaque monkey took a photograph of itself with it. When the photo was uploaded and shared, various parties began to argue over who held the copyright. Slater claimed it was his because it was his camera and he set up the situation. Others believed that if the monkey pressed the shutter, then the monkey, or no one, held the copyright.
The U.S. Copyright Office clarified its stance on the matter in the Compendium of U.S. Copyright Office Practices, Third Edition. It stated:
"The U.S. Copyright Office will not register works produced by nature, animals, or plants. Likewise, the Office cannot register a work purportedly created by divine or supernatural beings, although the Office may register a work where the application or the deposit copy state that the work was inspired by a divine spirit."
https://www.justice.gov/archives/jm/criminal-resource-manual...
The only difference is a machine doing it at a larger scale.
If you steal a book about how to make money flipping houses and then start flipping houses, at worst you are on the hook for a minor theft. The author can't sue you for illicitly learning to flip houses from the stolen book. That is just not how that works at all.
ChatGPT was trained by copying works into a dataset and using that for commercial purposes. Surely you can see how that's different.
I never had a book tell me to an accept a license agreement before I could read it.
It will say something like 'all rights reserved' and that you can't reproduce any part without permission, except for limited cases.
Please go get a book off of your shelf and look, I'm begging you.
Edit - Example from Infinite Jest: https://burnsiderarebooks.cdn.bibliopolis.com/pictures/14094...
If I read a math book that shows me how to do an integral, then use that knowledge to do integrals, I'm not infringing the copyright of the book, ffs.
If you read in a book that the main export of Germany is Bavarian creme doughnuts, and you use that knowledge in a job interview (i.e. making money) to land a job as a Bavarian creme doughnut importer, that is not copyright infringement. People learn things and put that knowledge to use, often for profit.
What kind of insane interpretation of copyright law are you working with?
A copyright notice is not an agreement. It's not a contract, it does not offer any consideration to the other party. There is no meeting of the minds. That is not at all how any of that works.
Inputting entire books verbatim into a model is not the same as a human learning a fact and using it later.
The model just trains a set of weights for a neural network, then probabilistically generates text in response to prompts.
And yeah, humans "input entire books verbatim" (aka, reading) and then regurgitate that knowledge later when they determine it is likely to be an appropriate response to a question or comment from someone else. It's really not that different. It doesn't even have to be a fact. How many responses on the Internet are just quotes from movies or TV shows?
We're not talking about copying verbatim, that's the whole point.
A human copied the entire text of many books - verbatim - into the training corpus.
Under copyright law, the human was free to read the whole thing. But not copy the whole thing into an AI model with the intent to profit.
Edit to add another example because someone is going to say that parody is its own thing and exempted. If I want to make money by writing a film in the same way Tarantino does I go and read all his screenplays to understand his style, pacing, character archetypes, etc. Should that also not be ok?
But, as with many things, the legality would ultimately depend on the nuances of the implementation.
I think it's different if OpenAI acquires licenses for all of the material it uses for training.
ML training doesn't work without having a copy of the data (books and screenplays). That data can either be copied in a non-infringing way (buy the books; acquire a license) or in an infringing way (download the books from a corpus without explicit permission like these AI companies are accused of doing). LLM services need a license just like Facebook, Instagram, etc. need a license to republish the stuff you post. The copyright holder maintains their copyright and the service that republishes their work (Facebook, Instagram, etc.) or publishes derivative works (AI) needs a license.
If you literally only consumed Tarantino media and had no other influences, and then one day made a movie so great everyone was calling you the next Tarantino - that would be fine!
As long as you didn't copy anything that he actually made himself.
Being influenced by is not the same as copying.
If it spits it out because it was in the training set or prompt, yes. If it spits it out and it was in the training set (or prompt), it is at least difficult to make the case that it is not infringement (if it was the prompt, then you basically have the same problem as any other case where you are proving that production of something that matches something someone else has a copyright on was independent, since copyright only protects against copying, not coincidence.) If it spits it out but it was not in the training set or prompt then clearly it was not infringement.
> But if it ingests copyrighted material and then spits out entirely new content influenced by the originals… how is that an issue?
Because a non-perfect mechanical copy is still a copy violating copyright if there is no license, and a derivative work produced by a human using AI as a tool is still an infringing derivative work if there is no license. While copyright protects against perfect copies, it protects against imperfect copies and derivative works as well. (Except where exceptions like Fair Use apply.)
1) Is ingestion of copyrighted material as part of the model training process an infringement itself? I say no - it is equivalent to a human going to a library and reading all the books there. The knowledge gained by the human is equivalent to the weights that end up in the model.
2) Is the output of copyrighted material from a model infringement? Yes, obviously. In the exact same way as a human regurgitating a duplicate or close-to-duplicate of someone else's book/song/script/whatever is. We already have an entire body of copyright law to cover this and don't need to reinvent the wheel just because we swapped a human creator for a machine.
It might be Fair Use for other reasons, but any argument that uses an “it’s like a human doing X” analogy is, legally, misguided. The courts simply do not see machines as being legally analogous to human brains. Humans can “copy" content into their brains through their senses and its not only not a copyright violation, its not legally a copy for which you need to do anything like fair use analysis, copy it, even lossily, into a machine where it is non-transiently stored, then it is, legally, a copy, and a violation unless an exception like Fair Use applies.
> Is the output of copyrighted material from a model infringement? Yes, obviously.
Again, only if it is actually a copy and not coincidence, and only if exceptions like Fair Use don't apply. It may be that training is more likely to be Fair Use than inference is—I’ve seen good arguments that don’t rely on treating machines as analogs of humans for that—but if you are copying copyright protected material, lossily or not, with a machine, it's going to be infringement unless an exception to copyright, like Fair Use, applies.
Even if OpenAI maliciously ingested copywritten work, it is just a bunch of numbers and if you go in and ask for it spit the book back out, it won't. That simply isn't how this works.
I dont understand the purpose of your feelings comment. I have no affiliations or preferences for any dogs in this fight.
I don't necessarily think they are identical scenarios but if I were OpenAI's lawyer that's probably what I'd try and point at.
I didnt say it would spit the book back out, at any point. I've said a human chose to copy the entirety of copyrighted texts, verbatim, into the training corpus for a language model that they intended to sell.
OpenAI isn't pirating any books content nor is it distributing or reproducing it. Nor are they even "stealing the book".
Both are transformative content.
But I understand what you're saying and it's certainly possible you're correct, I just don't think it's obvious.
Uh-oh, better tell the colleges and universities to stop promoting degree programs off the back of the potential increase in lifetime earnings. Wouldn't want anyone to get the idea that it's okay to absorb a bunch of copyrighted material during college to update their mental model, then make money by selling the output of that model for the rest of their lives.
Let alone the huge shift of power from one to many (one university to many students), to many to one (many resources to one company).
It seems to me OpenAI have done the equivalent of distilling a bunch of knowledge down into a book. And they’ve given that book a really good index.
They’re essentially an encyclopedia vendor. Just an exceedingly sophisticated one.
Back in the olden times publishers used to pay people to write encyclopedia entries based on summarizing stuff they had read in other books. Nowadays we still do the same thing only with volunteer time (Wikipedia requires everything it contains to be externally sourced, after all. You can’t write anything into a Wikipedia article without reading it somewhere copyrighted first).
OpenAI’s model isn’t so different. Source material, indexed and summarized to make it easier to search and use.
While the number of plot structures are fairly countable, a plot can still be considered different by changing part of its contents (character, setting, etc.). The decision on whether a plot outright infringes on an existing plot varies case-by-case. An example of outright copying is "Fistful of Dollars" directed by Sergio Leone, which lifts the plot from "Yojinbo" by Akira Kurosawa.
2. Romeo and Juliet
3. Robin Hood
4. Crime and Punishment
5. maybe there are only four.
I dunno, as long as you swap out genre tropes, make slight changes of geography, and/or change a few plot points, coolpying plot is routine in Hollywood.
the process of training requires reproduction and distribution of the works internally as part of the data processing pipeline so why wouldn't you need a license for that?
Again there is no "compression of information" in deep learning.
i would strongly disagree. when you are training a model you are taking the information from a document and extracting the relationships between tokens and storing that information conglomerated with the same information from a massive amount of other documents. the model that results is a compressed form of all of the information from all of the documents where you have extracted and stored a synthesis of the relationships between the tokens in all of them. this is a lossy compression, but it does reproduce exact sequences of source documents in some cases, so the original information is stored there.
you can very plausibly argue that an LLM model trained on copyrighted material violates the copyright on every single copyrighted document that was fed to it.
you would have to ask a judge about that because its a novel legal question. obviously nobody anticipated this technology at the time it was written so ultimately it will have to be a court that decides how to apply existing laws.
Original paper: https://huggingface.co/papers/2308.05374
Abstract:
major categories of LLM trustworthiness:
reliability, safety, fairness, resistance to misuse,
explainability and reasoning, adherence to social norms, and robustness.
Each major category is further divided into [...] a total of 29 sub-categories.
The measurement results indicate that, in general,
more aligned models tend to perform better in terms of overall trustworthiness
The paper shows that trustworthiness is now a design goal. It would seem good to avoid e.g., hallucinations, but is it really socially good to induce reliance?Compare AI companies targeting "trustworthiness" today with social media companies targeting "engagement" in decades past: they might increase uptake, but build significant externalities into their business model, and then lose most of the social benefits of adoption.
I read Harry Potter, am I now not allowed any magic spell names without paying a fee to JK Rowling for the training?
The AI grift continues. Copying the same Silicon Valley playbook with a new narrative for promoting their AI snake-oil.