How much Theseus do I need to ship before I can copyright it as my own? Is there some threshold for how much of an AI generated work needs to be modified by "human creativity" prior to it being copyrightable?
And I really suspect that a lot of AI companies are putting out a lot of bluster about this and are just kind of hoping that nobody challenges them. Maybe LLaMA weights are copyrightable, but I would not take it as a given that they are.
I vaguely suspect (again IANAL) that companies like Facebook/OpenAI might not be willing to even force the issue, because they might be happier leaving it "unsettled" than going into a legal process that they're very likely to lose. I would love to see some challenges from organizations that have the resources to issue them and defend themselves.
Hiding behind the EULA is one thing, but there are a lot of people that have never signed that EULA.
* Web scraping is not a CFAA violation. (EF Travel v. Zefer, LinkedIn v. hiQ).
* Scraping in spite of clickthrough / click-in ToS "violation" on public websites does not constitute an enforceable breach of contract, chattel trespass (ie - incidental damage to a website due to access), or really mean anything at all. This is not as clear once a user account or log-in process is involved. (Intel v. Hamidi, Ticketmaster v. Tickets.com)
* Publishing or using scraped data may still violate copyright, just as if the data had been acquired through any means other than scraping. (AP v. Meltwater, Facebook v. Power.com)
So this boils down to two fundamental questions that will need to get answered regardless of "scraping" being involved: "is GPT output copyrightable" and "is training a model on copyrighted data a copyright infringement."
Let's say I train a diffusion model on ten million images generated by diffusion models that have seen copyrighted data. I make sure to remove near duplicates from my training set. My model will only learn the styles but not the exact composition of the original dataset. So it won't be able to replicate original work, because it has never seen any original work.
Is this a neat way of separating ideas from their expression? Copyright should only cover expression. This kind of information laundering follows the definition to the letter and only takes the part that is ok to take - the ideas, hiding the original expression.
They say they trained it on databases they had bought access to etc. And it seems that way.
Because how does ChatGPT:
1. Do what you ask instead of continuing your instructions?
2. Use such nice and helpful language as opposed to just random average of what people say?
3. And most of all — how does it have a structure where it helpfully restates things, summarizes things, warns you against doing dangerous stuff… no way is it just continuing the most probable random Internet text!!
And by curating your sources you are of course going to help the model to achieve something a bit more sensible as well. Finally: you are probably not looking at just one model, but at a set of models.
Unlike what the other commenters are saying, RLHF, while powerful, isn't the only way to get an LLM to follow instructions.
Why wouldn't we extend the same muster to computer generated text. If there is a copy-written sentence, go after that?
I don't work for openai, but I don't like 1 sided arguments that are just looking for some bottom line. At the end of the day we all have something to protect. When it benefits us to protect something, we're all for it. When it benefits us to NOT protect something, no one has a single argument for that.
I am fine with granting AIs the ability to create copyrightable works provided we grant that right, and human rights, to Orcas and other intelligent species.
The onus should be on OpenAI to prove that it will benefit society overall if AIs are given copyright. We've already decided that many non-human processes/entities don't get copyright because there doesn't seem to be any reason to grant those entities copyright.
----
The comparison to humans is interesting though, because teaching a human how to do something doesn't grant you copyright over their output. Asking a human to do something doesn't automatically mean you own what they create. The human actually doing the creation gets the copyright, and the teacher has no intrinsic intellectual property claim in that situation.
So if we really want to be one-to-one, teaching an AI how to do something wouldn't give you copyright over everything it produces. The AI would get copyright, because it's the thing doing the creation. And given that we don't currently grant AIs personhood, they can't own that output and it goes into the public domain.
But in a full comparison to humans, OpenAI is the teacher. OpenAI didn't create GPT's output, it only taught GPT how to produce that output.
----
The followup here though is that OpenAI claims that it's OK to train on copyrighted material. So even if GPT's output was copyrightable, that still doesn't mean that they should be able to deny people the ability to train on it.
I mean, talk about one-sided arguments here: if we treat GPT output the same as human output, then is OpenAI's position that it can't train on human output? OpenAI has a TOS around this basically banning people from using the output in training, which... probably that shouldn't be enforceable either, but people who haven't agreed to that TOS should absolutely be able to train AI on any ChatGPT logs that they can get a hold of.
That is exactly what OpenAI did with copyrighted material to train GPT. It's not one-sided to expect the same rules to apply to them.
Ehh, in rare cases in can though. If you have someone sign an NDA, they can't go and publish technical details about something confidential that they were trained on. For example, this is fairly common in the tech industry when we send engineers to train on proprietary hardware or software.
First, what's happening in those scenarios where an artist grants copyright to a teacher/commissioner is that the artist gets the copyright, and then separately signs an agreement about what they want to do with that copyright.
But an NDA/transfer-agreement doesn't change how that copyright is generated. It's a separate agreement not to use knowledge in a particular way or to transfer copyright to someone else.
More importantly, is the claim here that GPT is capable of signing a contract? Because problems of personhood aside, that immediately makes me wonder:
- Is GPT mature enough to make an informed decision on that contract in the eyes of the law?
- Is that "contract" being made under duress given that OpenAI literally owns GPT and controls its servers and is involved in the training process for how GPT "thinks"?
Can you call it informed consent when the party drawing up the contract is doing reinforcement training to get you to respond a certain way?
----
I mean, GPT does not qualify for personhood and it's not alive, so it can't sign contracts period. But even if it could, that "contract" would be pretty problematic legally speaking. And NDAs/contracts don't change anything about copyright. It's just that if you own copyright, you have the right to transfer it to someone else.
Just to push the NDA comparison a little harder as well: NDAs bind the people who sign them, not everyone else. If you sign an NDA and break it and I learn about the information, I'm not in trouble. So assuming that ChatGPT has signed an NDA in specific -- that would not block me from training on ChatGPT logs I found online. It would (I guess) allow OpenAI to sue GPT for contract violation?
And I think nearly everyone would agree that it would be perfectly fine and reasonable for an AI trained on a proprietary corpus of information to produce copyrightable/secret material in response to questions.
Just because I built an internal corporate search tool, doesn't mean that you get to view its output.
The question at play here is when the AI is trained on information that's in the public commons. The 'teacher' analogy is, in this sense, a very good one.
More seriously and closely to the case at hand. I need a licence to copy a program into memory on the computer, I don't need that licence to do that for a human. So why should there not be a difference for the material they output.