I don’t know which view is better since I’ve only just realized there are more than one.
I don’t know which view is better since I’ve only just realized there are more than one.
That's not how copyright or patent law works.
If I publish something on the internet, I am only making it available for people to read.
That does not give anyone the right to use it in other ways.
https://www.copyright.gov/help/faq/faq-fairuse.html
We make use of this exception on HN all the time, like this, from the link above:
> How much of someone else's work can I use without getting permission?
> Under the fair use doctrine of the U.S. copyright statute, it is permissible to use limited portions of a work including quotes, for purposes such as commentary, criticism, news reporting, and scholarly reports. There are no legal rules permitting the use of a specific number of words, a certain number of musical notes, or percentage of a work. Whether a particular use qualifies as fair use depends on all the circumstances.
It makes sense there is no limit on the number of words that can be used under fair use, but it's certainly less than all of them.
The questions around LLMs learning from copyrighted material are still open and need to be settled in court. I personally imagine finding infringement would impose more harm on society and progress than letting the models acquire knowledge from these copyrighted works.
I'm gagging at the nonsensical anthropomorphizing being done to end-run the fact that what the LLMs are doing is copying.
Please chill
> the fact that what the LLMs are doing is copying.
I disagree, the training process creates token representations and weighted connections between them. The models later produce probabilistic token sequences, not so unlike what our meat bodies do, though by very different mechanisms. The fact that certain sequences can be reproduced verbatim is likely a consequence of overfitting. They certainly cannot reproduce all training data verbatim. It would be interesting to know the features around what can and cannot be, and how.
Your response to me calling out your baseless anthropomorphizing was to double down on it? It's amusing to me that you don't think you are condescending.
ChatGPT can help you with that ;]
If the AI models are reproducing copyrighted works, then that's a problem. And it does look like there are some examples where that might be happening beyond notions of fair use. But slupring up copyrighted content to train a model seems to fall under allowed use.
I just started writing a new novel. It's an interesting, in my opinion highly novel fantasy/SF(ish) story, for once not fanfiction of anything that's still in copyright -- most people wouldn't count stories based on ancient norse mythology as 'fanfiction' -- but that doesn't mean it isn't derivative. It means, instead of naming two or three things it's derivative of, I can name ten to fifteen.
That's normal. All stories are derivative, and if you point me at an author who claims theirs aren't, you're pointing at a liar. The job of an author is to put the building blocks together in a new and interesting form, not to make them up from whole cloth. It's impossible to invent more than two or three truly novel ideas per day, even if you're incredibly imaginative, and most of those won't be any good.
The difference between humans and AIs, nowadays, seem to be that the AIs use millions of sources instead of ten to fifteen. Or, alternately, that they use none -- and theirs is less derivative -- because certainly everything I've ever read goes into my writing, not just the things I recognise I'm using.
No. Full stop. Humans aren't stochastic parrots. Pointing to a lack of understanding about what exactly happens in the human mind is, FULL STOP, not evidence that LLMs are doing the same things humans do.
Humans are not stochastic, they're obviously chaotic[1]. Which is to say: not parrots at all.
Some of the modern models I've seen also seem to be chaotic too though, so that's interesting [2]. I'm going to assume LLMs probably exhibit the same properties.
[1] https://en.wikipedia.org/wiki/Chaos_theory (Chaotic systems sometimes seem to be stochastic, but they're actually much stranger and more interesting!)
[2] I've been messing with stable diffusion to get a feel for (and/or avoid) tipping points: that is to say, points in latent space where the model becomes very sensitive to small changes in initial parameters. You can find instances fairly quickly even by hand by doing bisect search.
That's quite an assumption to make.
That's not my argument. My argument is that the anti-AI arguments, as spoken, also match to what I know I'm doing as a human. In my opinion better than it matches to what the AIs are doing, because as you say, they aren't human.
As for fair use, this is not the same as someone remixing or sampling songs, or writing fan fiction or satire or quoting works for criticism.
People do not want models trained on their creative works, so that someone else can make money using those models to produce similar creative works as a service for third parties.
It in no way blocks right to absorb, understand, or stand upon for the next idea.
There is copyright law, and there is patent law, and what's true for one is almost never true for the other. Copyright is automatic. Patents are applied for and cost money. Copyright lasts 70 years from your death. Patent's last 25 years from invention. Copyrights don't prevent other people doing the same thing you did. Patents make it illegal to do the same thing, even if you separately invented it and didn't know about the first one etc.
Publishing the work, or even registering it with the Copyright Office, are optional steps that might make it easier to prove your claim if there is a dispute, but that's conceptually separate from your actual rights.
You have the right to control who can copy and redistribute your work, which is a totally different ball game from all of these models.
Plus, change copyright, and it doesn't fix a thing except for preventing everything from being open source, because groups like Facebook and Google and all the other websites have you agreed to a terms of service where you say that you agree for the websites right to use the image however they want.
You kind of glossed over the idea that, no, it's not a totally different ball game without really providing any explanation of why.
Also, you have the right to control who can profit from derivative works of yours, which you didn't mention.
When you have StableDiffusion's creators being sued by Getty because they CLEARLY trained the model on Getty's catalog, which you can tell because some queries output a warped version of the Getty logo and it reproduces in obvious ways the original photo, there's some problems here. You've basically just made these ML models copyright laundering systems.
> Plus, change copyright, and it doesn't fix a thing except for preventing everything from being open source, because groups like Facebook and Google and all the other websites have you agreed to a terms of service where you say that you agree for the websites right to use the image however they want.
Maybe copyright is a terrible idea but it's what we're using today.
Because of I buy a book from a store, you now no longer control that book. I can read it. I can use it for just about any purpose I want to.
If you put a image up online free to download, it's much the same situation. I can download and use that image however I see fit.
I can't share that book to others.
I can't post that image online (in theory, in practice this is very common).
Copyright protects distribution, not use, because protecting use would be hideously draconian. To protect use means that the government can enter your home and validate how you use the stuff you've acquired. It says that you not only control your works, but all things even remotely derived from it.
Derived works <> preventing machine learning model translation. A derived work is typically something like adding a chapter or translating.
This would kill video game genres.
This would kill novel genres.
This would kill most of the internet and the ability to download images at all.
Open this door and the entire system of copyright will turn into an absolute monster. Every large copyright holder will use the precedent and they will argue with validity that they have the right to control your use of their material as well.
Adding that the model is violating copyright by responding images is perfectly valid, but this isn't what Getty and others seems to be aiming for. The want the right to charge to use to train, which is a gross extension of copyright that has never existed.
The answer here should be three fold.
Light a fire under stable diffusion to get their model to never overfit and clone or close to clone images.
Ensure people who distribute images that are clones made with the models can be hit with copyright.
Protect the ability to train models on publicly available content.
Protect people's right to distribute their content while not annihilating AI in the process. People making these arguments don't care about copyright, in almost all cases they care about either making money (getty) or annihilating AI (artists)
The implication here is that training ML models on your source material and then having it regurgitate that source material verbatim is not distribution.
> Derived works <> preventing machine learning model translation. A derived work is typically something like adding a chapter or translating.
The output of the model can be a derived work though.
> This would kill video game genres. > This would kill novel genres. > This would kill most of the internet and the ability to download images at all.
How? This is hyperbole. You're arguing that the ML model doesn't do "distribution" and are attacking "use" when that's not the real point.
> Open this door and the entire system of copyright will turn into an absolute monster. Every large copyright holder will use the precedent and they will argue with validity that they have the right to control your use of their material as well.
The DMCA already exists. I agree it's bad but we're taking law here not morality.
> Ensure people who distribute images that are clones made with the models can be hit with copyright.
How do you in good faith argue that ChatGPT for example is not doing the distribution? OpenAI are responsible for generating and showing you the output based on some instructions they have provided.
I would reread that part.
Models output very similar work to training data at times, even identical in some cases, but the vast majority of it is different enough to not be derived work.
The DCMA act effects public uses of art like YouTube. This sort of copyright where you can regulate how it's used does not exist in any form currently.
You can't make copies and sell them. You can't copy it into your own work. You can't create derivative works, etc. You can't use it in a movie. You can't use it to write lyrics. I could go on.
I can also*do* all of those things, I just can't sell it.
That's what copyright is.
Copy right.
Not "use for training an AI model" right.
It's not similar to using it in a movie or a "derivative work" either. Those almost always actually contain the work.
The only valid argument here are times than an AI model does reproduce the original, otherwise copyright has no domain of the use of images for training, unless you illegally downloaded the images, which you didn't if you used LAION, since it's all just public image links.
For example, lets say a not-awake-enough-in-the-morning admin at Equifax misconfigures a proxy server, allowing worldwide (ro?) access to what should be (only) internally viewable customer PII.
Clearly it's the "fault" of the company (Equifax in this hypothetical example).
But how to unwind the training of that PII data into the "AI", as that data cannot be allowed to "remain" in the AI system and be potentially regurgitated with the right prompting?