Do you think that DeepMind, OpenAI, etc. don't have this dataset copied 10 times over? Think again!
All the current LLMs are trained on data scraped from random websites, with no regard for the website's copyright, aren't they?
Presumably because these big businesses have a theory training an LLM is 'fair use'.
Why wouldn't they treat books the same way they treat web pages?
But it also feels like standing at the edge of the sea complaining about the tide coming in...I'm not sure there's really much that can / will be done about it.
That being said, it feels like there's also a shade of perspective from the old quote:
"In its majestic equality, the law forbids rich and poor alike to sleep under bridges, beg in the streets and steal loaves of bread." - assuming everything in public is fair game, then everyone is welcome to build a multi-petabyte database of text and use millions of dollars worth of GPUs to train an AI on it.
But saying "we steal to stay relevant, because we don't have as much funds as our competitors" is not the answer.
I understand that in Japan, for example, it is legal to apply deep analysis techniques, including deep learning and generative models to any data, regardless the copyright.
I would guess that in Japan, they’ve looked at China, where copyrights are also not an issue. And decided that it’s a really bad idea to stop research because of copyrights and let China to be first at the AGI race.
If a lot of time will get wasted on the GDPR and copyrights discussions in Europe and United States, instead of actually doing research, the democratic world will get behind.
Giant corporations that are scrambling to keep R&D going are quite law-abiding actually. And are currently struggling, as most of the datasets out there can’t be used, due to unclear copyright and licensing regulations.
Copyright ethics are something we have invented.
China chooses not to follow western copyright regulations. That’s not unethical, to them anyway.
Indeed, Copyright law was originally intended to allow people to legally copy things, not protect corporate interests for an eternity!
TLDR Model training will go "dark" or underground, with models distributed like illicit content.
The only protection is specific jurisdictions ruling this stuff is fair use, which is risky when you could be found liable in literally any country.
How are you supposed to know if that's because it ingested the text of Harry Potter or a bunch of fan blogs talking about Harry Potter?
> The only protection is specific jurisdictions ruling this stuff is fair use, which is risky when you could be found liable in literally any country.
Saudi Arabia has quite strict blasphemy laws but people don't seem to be bothered much by it when hosting quite obvious violations of them on their servers in North America.
Doesn’t actually matter here because those fan blogs are also derivative works. I may have read “Call me Ishmael” in a blog, but the quote is from Moby Dick.
ChatGPT can argue for a fair use exception, but combining lots of derivative works can be copyright infringement even if none of the things you directly copied where. IE: If you copy a lot of excerpts from a poem and recreate the poem that doesn’t mean you can now use the poem.
Being able to say Harry Potter isn’t directly in their training data is at best useful to arguing something was unintentional copyright infringement rather than not actually being copyright infringement.
> Blasphemy laws
That’s a criminal not civil issue and there aren’t a huge number of treaties on the subject. If a US company is sued in Australia for copyright infringement they don’t get to move the case to the US.
Not necessarily. To be a derivative work it has to contain copyrightable elements, not just a small number of words in the same order.
> ChatGPT can argue for a fair use exception, but combining lots of derivative works can be copyright infringement even if none of the things you directly copied where. IE: If you copy a lot of excerpts from a poem and recreate the poem that doesn’t mean you can now use the poem.
This is the argument for why the model itself isn't infringing, isn't it? All of the excerpts as presented in their original context isn't the same thing as all of the excerpts purposefully rearranged back into the original work.
It's like on the one hand arguing that a dictionary is infringing because it contains all of the words in Harry Potter, and on the other hand claiming that you can reproduce the text of Harry Potter without liability because you reconstructed it by rearranging the words in a dictionary. One of these things is not like the other.
> If a US company is sued in Australia for copyright infringement they don’t get to move the case to the US.
If a US company has no operations in Australia, how does an Australian court have jurisdiction?
Fictional character names are copyrightable elements. Sure, it’s no issue for actual place names, but those aren’t the words people will be looking for.
> a dictionary is infringing because it contains all the words in Harry Potter
The argument is the model may have recreated the story of Harry Potter from being fed so much information about it.
DALL·E can’t contain a copy of every image in it’s corpus simply based on information theory, the model just isn’t that large. However it has no problem recreating the Mona Lisa because so many of its source images included it. That’s no problem for something out of copyright but it’s a big problem for major works in our culture like Harry Potter or Star Wars where potentially hundreds of thousands of derivative works get stitched together into a copy of the story.
If ask ChatGPT to summarize Harry’s first year at Hogwarts and it spits back a detailed summary of his year including dialog etc that’s wildly past the de minimis standard. I assume OpenAI has tried to prevent this, but this kind of reconstruction seems like a legal landline.
Copyright, or trademark?
> The argument is the model may have recreated the story of Harry Potter from being fed so much information about it.
It may contain enough information to recreate it, but is that the same thing?
Suppose you're a major blog platform and everybody's blog contains the data the model was trained on, i.e. enough to recreate the story of Harry Potter if you were to read them all and then be asked to summarize it or write fan fiction. If your company isn't infringing the book's copyright, why would the model be? It certainly contains no more information than it was trained on.
The output might be infringing, in the same way as a person who read all the blogs to learn the story could write something that was. But the output isn't the model.
Always copyright, sometimes both. Really though fictional works have a really low bar for copyright infringement.
> It may contain enough information to recreate it, but is that the same thing?
Yes. The law currently doesn’t give a shit about how you got the information just it’s ultimate origin. It’s the same reason MP3’s absolutely qualify as as copies even though none of the bits map 1:1 with the source material.
An LLM that reproduces a lossy copy of some work is just another lossy format. A judge isn’t going to care about AI magic, they are just going to see the infringement. The only defense is to not create copies of copyrighted works which shouldn’t be an issue if this stuff is actually transformative.
Remember lawsuits aren’t beyond a reasonable doubt the standard is much lower.
The value of names to copyright holders is simply how distinct they are. John Smith isn’t particularly unusual but if you name a bounty hunter Boba Fett it’s hard to argue originality. Complete originality isn’t required as long as you’re safely inside a fair use exception, but overwhelming originality is required.
So it’s generally safer for a character to talk about their favorite parts of Star Wars than have your stories plot be a group of rebels trying to destroy a planet killing doomsday weapon guarded by a sword wielding space wizards.
My point is that the meaning of "copying" as described here is unclear unclear if it's not restricted to verbatim plagiarism. As is "elements", for that matter.
Without details of what specific data a language model was trained on, there is no way to differentiate it having retrieved that name from a review of the book that was used with permission from the author (which is their own copyright and fair use on its own) or from the books itself (which would be copyright infringement).
If those records don't exist, the only mechanism you have left is to try and get the language model to spit out sufficient verbatim text from the source material to cross the fair use threshold. This doesn't work in the case of "obvious extrapolation" either which is a whole other defense that could be used depending on what that body produced was (If knowing only that the main character is named Harry, who is a wizard going to a wizarding school in the British countryside is required to produce a couple of close looking paragraphs, you need a much closer match with the original text over those two paragraphs for it to be infringing).
I still would like to see if legally a judge would consider the training illegal. I do not think this got tested and i would not consider the answer as clear cut. If a company buys a million ebooks, can it simply train an AI on it? If no employee reads the books and only the AI "reads" it you could argue that the company bought it only for a single reader/entity. They will probably need to update the copyright system to take this case into account. We will see in a few years how this plays out.
In the end these authors will eventually lose their copyright after their death. And nothing will stop the AI training at that point. This only delays the inevitable by a few years. The public domain already has a lot of amazing books and resources. Google even has access to very rare books nobody else has access to digitally i think:
https://blog.google/outreach-initiatives/arts-culture/in-beg...