Used? Yes. Reproduced? I don't think that necessarily follows.
> I don’t know, but to claim that they’re just storing a url is bogus.
They're definitely not just storing an URL, but claiming that what they're storing is some kind of "copy" seems equally bogus.
Where have they reproduced his content?
You don't need a license for your fair use of something. Can you walk me through how downloading a photo to your personal computer and studying it, perhaps to become a better artist is copyright infringement? Now let's apply that analogy to training a model, is the process of using a particular image to train a model, ultimately reconfiguring a bunch of weights in a matrix, so it "learns" whatever particulars it learns, copyright infringement? Is training a painting monkey that may paint something similar to a copyrighted painting, on a copyrighted painting, copyright infringement?
I personally do believe that training using copyrighted data is _not_ a violation of any existing copyright, but using the trained AI to reproduce and distribute something that an expert in the field would regard as a copy is a violation.
FWIW I do believe training a model is fundamentally different from training a human.
The obvious analog is facial recognition, deemed privacy violating by laws and courts in multiple jurisdictions, vs a human doing the same thing manually.
It was deemed "privacy violating" under specific privacy laws within those jurisdictions.
i.e. go pass a law amending copyright law if you think so. there's no substantial reproduction here, in the context of stable diffusion, until there is, and then yeah in that instance it's probably a copyright violation, like a substantial reproduction made by a human would be. but training itself, probably not. but yes, there have been no rulings on it in this context.
> go pass a law amending copyright law if you think so
I don't need to do anything. Let's see until one of these cases come to the courts. You're talking as though it's a foregone conclusion that they'll rule the way you think. I wouldn't be so sure myself.
Until we pass 21st century laws, thats what applies.
Do you have a link to a court decision to that effect?
How, specifically, is "training an AI" different from "training a person" here?
Training(2) software is to training(1) a human what firewall is to an actual wall made of fire, or what killing a process is to killing a human. The word is a synonym—more precisely, industry slang that happens to resemble a word used outside of the industry due to superficial similarity.
So if you wonder how training(2) software is different from training(1) a person, first understand that it’s a different word that never signified the same concept.
Subsequently, if you want to argue that training(1) should apply to software now, the onus would be on you to explain why and how so. You should be prepared to argue that software is sufficiently like a human—you know, that it understands ideas, has agency and free will, thinks like a human and has human-like consciousness, is capable not only of performing instructions given to it by its operator and act as a tool but to actually consider its actions and make own moral judgements, things like that.
And if you come up with satisfactory evidence in favor of all that, and have grounds to believe some software is enough like a human that training(1) applies to it, then why are you fighting to allow its operator to ignore copyright—and not for more important things, such as to free this human-like being from abuse by its operator and grant it basic human rights? If we imagine that software is like a person enough that “training” it means the same thing as training a human, we should be prepared to acknowledge all the implications that come with that.
The reality is that there are companies who would like us to both believe their software is human-like (so that we don’t sue them for rampant copyright abuse) but also not at all human (so that we don’t demand them to stop profiting from what would be slave labor). Naturally, if they pick one or the other they stand to lose from “a lot of money” to “entire business model”—but we should help them make that choice.
How does that change the fact that there is no substantial reproduction of the works in the resulting weights, or in the "output" of these models with the weights?
Can you point to me where it's illegal to take copyrighted works and "train" neural networks on them, as long as there's no substantial reproduction of the works in the output of that process, or in the output of a particular configuration of "trained" and "frozen weights"?
However you define it, training typically doesn't involve making permanent copies of the data.
This means that Under EU law -as far as I can tell- it is probably legal. Under US law the different circuits have a slightly different interpretations of the law, but probably would agree that this is fair use.
<Training> by any definition you care to give it does not rise to the level of copyright violation (In the case of training eg stable diffusion, on average only a couple of bits worth of "notes" are stored per image. If that's a copyright violation then pretty much anything would be). The main issue -believe it or not- is actually the part where images get temporarily downloaded.
As stated: explicitly legal in the EU. May need some work in the US.
You are talking about two different words, again. Human learning is not a copyright violation. Machine learning is. Machine is not human.
> The main issue -believe it or not- is actually the part where images get temporarily downloaded.
No, the main issue is using a tool to sell derivative works automatically generated from copyrighted original works. If images/text is not stored it wouldn’t change a thing.
If images/text being stored or not is not part of the argument, then what IS the argument that AI works are derivative? Is it only certain works? All works? Is it automatic? Does it require human intent? How do you get from A to B here?
And of course there is human intent, what are you even talking about? This is law. Law is sort of centered around human actions and intent plays a big role. In this case, operators fully intended to scrape copyrighted works, feed it to this tool and operate it for profit (because money smells good).
“Our client fundamentally understands that your client may not like the temporary reproduction of his works,” LAIOn lawyers wrote to Kneschke’s. “Only this has been expressly permitted by the European legislator.”
Which is exactly what happened with stable diffusion - say what you will about it, it’s distributing the ability to create value from AI among a larger number of benefactors than OpenAI
People don't seem to be getting quite yet what 'artificial intelligence' intrinsically really means.
But they will.
If the artist doesn’t want people and algorithms to see the image, then put a password in front of it. Or add no-index to robots.txt.
This is like driving around with a sign on your car and complaining that people are looking at it and imagining the image in ways you don’t like. If you don’t want people to see the image, don’t put it on a sign and drive around.
There's a bit of nuance in this case, but in general, yes, you need a licence.
There is no law that says you cannot train AI models on copyrighted data.
Also, It is legal to make copies of copyrighted materials so long as you don’t distribute them.
Also you do not have to preserve the original for copyright to apply: any extract is only allowed under well defined fair use rules. A GAN's purpose is definitely not parody, commentary, ...
That is not a definition that people use, nor the legal system
> Taking data and putting it into any form of storage, ie weights in a model, is called copying. Producing an output out of the gathered data is then a step in distribution.
It isn’t though.
If I train on Batman and I produce Spider-Man, the Batman copyright holder cannot sue me. It would be dumb to suggest otherwise.
This will see its day in court, and clearly indeed the meaning of the verb "train" will feature prominently in the debate.
I tend to side with reliance on dictionaries for the meaning of words, and in my understanding the courts also do.
If you load up only batman pictures on your model you will clearly only produce batman pictures. No one is being dumb here that does not want to be, I am sorry.
Perhaps. Or, more likely I think, your argument is weak and doesn’t contribute meaningfully to the discussion.
Personally, I downvoted you do to my decades old rule of “downvote messages that complain about downvoting.” I don’t always downvote whiny comments, but I always downvote whiny comments about downvotes. (and I will be downvoting my own message as I am both pedantic and reliable).
That’s not what I’m claiming though. I’m saying I take a model that has been trained on Batman and other things and I produce Spider-Man. If your point of view is to believed, the Batman copyright holder can sue me for producing a work that contains Spider-Man.
That’s definitely how copyright works.
I don’t think training on these has been tested, but I expect it will be allowed, otherwise how can a web browser receive, cache, and render a copyrighted photo? I think the closest is there were some lawsuits about google caching websites that contain copyrighted material. I assume they worked out in Google’s favor as google (and others) cache web sites even though they contain copyrighted material. The exception I remember is there are some specific rulings and laws governing how news is treated.
"If the image is available publicly, without restriction" — is irrelevant.
Why quickly skip past the most unjustifiable part? Because you're about to use a human metaphor (like "see") for an algorithm. My algorithm is "copy." So instead of driving, it's like selling a book with a sign on the book that says "copyright 2017, all rights reserved" and saying that people can't copy it. If you don't want people to copy it, don't write and distribute the book.
I actually believe this, but I don't know how somebody believes in one algo and not the other and remains consistent.
LAION doesn't do the SD training themselves though. This just provide links with annotations.
[1] Exception: The Ephemeral copy that is made when actually downloading the image. This is permitted by EU law. AFAIK after the image has been analyzed for certain properties like "red" or "traffic light" or "bus" or "vertical line" .. .etc... it is then discarded.
[1] https://lawreview.law.ucdavis.edu/issues/53/5/notes/files/53... with thanks to https://news.ycombinator.com/item?id=35713309
[2] https://eur-lex.europa.eu/LexUriServ/LexUriServ.do?uri=CELEX...
[3] https://www.gesetze-im-internet.de/englisch_urhg/englisch_ur...
"""Article 5
Exceptions and limitations
1. Temporary acts of reproduction referred to in Article 2, which are transient or incidental [and] an integral and essential part of a technological process and whose sole purpose is to enable:
(a) a transmission in a network between third parties by an intermediary, or
(b) a lawful use of a work or other subject-matter to be made, and which have no independent economic significance, shall be exempted from the reproduction right provided for in Article 2. """
As upload to a GAN is not transmission, and the economic significance is pretty clear.
Then the German 44a is an exact copy of the EC directive, so this also applies I think.
As the word "train" refers to applying the uniquely human capability of acquiring information, the word "upload" seems the most appropriate term fallback to use when referring to putting information into some AI system.
You're welcome!
Reference: open any dictionary and look up the word "train". Choo choo!
"Upload" -in my mind- is typically a network operation (and in the opposite direction at that) , which is not what is happening here. I would prefer to stick to terminology used in the sources and literature, if possible (as opposed to inventing a new nomenclature on the spot).
A very brief summary/introduction of the operations involved can be found in section I. A. of [1]. (a source I referred to previously)
[1] https://lawreview.law.ucdavis.edu/issues/53/5/notes/files/53...
Dogs are trained to do certain things, are dogs humans?
In the sense that yes, the bits are copied from one register to another. But they aren’t stored permanently, nor are they distributed. So there’s no copyright violation.
Similarly my brain is making copies when I look at something. And routers on the internet make many copies sending the material to me. Etc etc.
If training on an image is copyright infringement for making copies, then so is viewing the image through glasses (image is transformed by the lens) or dragging the image from one monitor to another or caching the image in your browser and viewing it later from the cache rather than retrieving it again.