Artists put content on the internet so you can look at it and share it with your friends and say “wow cool art” not so you can run it through an art generating model to put them out of job. That we never anticipated this possibility means you call fallback on the “fair use” argument but legal does not mean ethical, especially when the law has yet to catch up to changing times.
But they are not simply sharing a single link and the “intent” is different. I personally feel that we should not treat human learning and AI learning the same. Unlike humans, the AI is essentially immortal and infinitely copiable.
[0] The answer at least in the US is the act itself is likely not illegal although it can be construed to lead to a crime, such as incitement.
Now, let's try something similar. You have a monkey you want to train to paint. You download an image of a painting and give it to the monkey to train it on, amongst many. The monkey, in the future, may paint something that has vaguely or substantially similar qualities to the original painting. Has copyright infringement occurred in this instance?
Let's go further: You have a neural network you want to train to paint. You download an image of a painting and give it to the neural network to train it on, amongst many. The neural network, in the future, may paint something that has vaguely or substantially similar qualities to the original painting. Has copyright infringement occurred when that happens?
I find it hilarious that in this thread, instead of asking the question of whether this is fair use, multiple posters have jumped to comparing this to thievery, promoting terrorism and posting of private information as an analogy to harm.
https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...
The neural network behind the LLM is "trained." Yes, it's a term of art, very observant. That doesn't change the fact that there's not a substantial reproduction of a work in it. LLMs "learn" to predict text by feeding them vectors, the resulting "weights" of that process could be thought of and implemented as a big walkable decision tree to predict the next token it weren't for combinatorial explosion, but there's no substantial reproduction, or "copy", in there.
Is there a substantial reproduction, in your head, of a song or a text or a painting you can perfectly recall and can't seem to forget? Be careful, the copyright police will get you, the only way to delete it is to delete yourself!
>feeding into a machine is called copying
Copyright law deals in "reproductions" as far as I understand it. Feeding a copyrighted document as vectors into a program that reconfigures some matrix of vectors, and then using that matrices to output probability distributions into something tangible, where something tangible isn't a complete or substantial reproduction of the copyrighted work, is not a copyright violation.
If you think it is, amend copyright law, or take it to court and let a judge settle the matter, and hope it's in your favor.
About your question: """Is there a substantial reproduction, in your head, of a song or a text or a painting you can perfectly recall and can't seem to forget? Be careful, the copyright police will get you, the only way to delete it is to delete yourself!"""
This equivalence between "perfect" human recall, and a copy of the data input into an AI model is a bit of a strawman I think: it is the distribution of copied information that copyright protects against, the information has been uploaded in vectors, have fun demonstrating it is not being used in producing this or that model output. If I learn a popular song and do a cover, there is a copyright law for that, I owe rights.
The "vectors" are not "uploaded" or "copied" into a file or neural network, they're transformed. In the context of stable diffusion, They're transformed, progressively "noised" or "corrupted", "diffused" with random gaussian noise, and in the context of stable diffusion, it's "trained" and "learns" how to "denoise" various "noised" stages of images represented as a vector of pixel data into their original form.
Then, when it comes to generation of images consistent with an "annotation" or "prompt", it is "conditioned towards" or "biased towards", with more "training", by noising an input image, and concating or combining that vector of pixels with a vector of the annotation of the image. It then "learns" to denoise with that conditioning information, the annotation.
Then, you can take the trained model, and do the same thing, with just a text prompt as a vector concated to a vector of random gaussian noise, and no input images.
That's basically and very simplistically how it works.
The output is not a substantial reproduction from the input images + annotations when trained. It takes the random noise, and "tries" to denoise it into something consistent with the prompt with conditioning to guide it.
Your attempt at covering would be a substantially similar reproduction. Your goal is to do a reproduction. Whereas, the model "learns" to generate images consistent with an annotation/prompt, by conditioning it with that "goal" on top of how it "learned" to denoise the images.
Selling pipes, cellphones, igniters, and black powder is perfectly legal. Selling pre assembled pipe bombs filled with black powder and a remote detonators is going to very quickly get you into a great deal of legal trouble.
2. The artist made the photo publicly available.
3. Read this opinion and come back to the conversation with a more level head instead of making analogies to promoting terrorism: https://lawreview.law.ucdavis.edu/issues/53/5/notes/files/53...
2. The artist disabled crawling on their website so it’s accessible but not fully available.
3. Your link says many things that aren’t settled in terms of US law but rather the options of the author. “even if infringement occurs during machine learning, training AI with copyrighted works would likely be excused by the fair use doctrine.” That’s a lot of possibilities not actual guarantees.
However, it ignores the largest issue namely if the output of these AI’s are all simply derivative works.
3. That's why it's called an opinion. No shit. Maybe instead of talking about promoting terrorism, you should have a level head and talk about fair use.
Something people don’t understand is “Right click file save as…” on a copyrighted image breaks copyright, as you don’t have permission to make a copy. The do have implied permission to make incidental copies to view a website, but that’s it.
In all seriousness, that's total shit. It's publicly accessible. That's the permission. There have been different cases like weev where the judge has viewed it differently because of specific details like carrying out the action of enumeration, or guessing IDs, to reveal non-public information, but that had to do with the CFAA.
If you reproduce content in total in another work, resell, etc., that's another thing.
Fair use is permission. Every library book is accessible that doesn’t have copyright implications.
In the context you're now talking about by the letter of copyright law, downloading a photo, image, file of something or other, a "copyrighted work" where it's publicly accessible, and has a particular license specified that you may not reproduce it, may technically be "unlawful" letter by letter of the law, but I doubt any judge is going to actually see it that way versus intent of the law unless you literally reproduce or share a substantial portion of it, sell a complete copy, etc. It's almost certainly fair use to study, and use portions of the copyrighted work in your own copyrighted works.
As to making a copy with intent to X, that’s what fair use is. A student may photocopy the full text of an short article so they can accurate quote it in their term paper. They can’t simply photocopy an article because they are a student with access to the article and a photocopier nor can the photocopy a full book because they want to use a short quote in their paper. Copying incidental to acceptable use becomes retroactively acceptable. This distinction may seem crazy to you, but the fact that intent matters means you can’t judge an action without context.
Fair use in commercial context is looked at with vastly more suspicion than fair use in academic context, which again demonstrates specific actions on their own aren’t always enough information to say if something is allowed.
It's been determined that it's legal, if it's publicly accessible, and you don't receive a cease-and-desist letter telling you to stop. If it's public, that's permission. Simply "turning off scraping" by putting a robots.txt there doesn't make the content and linked images of a public web page any less public and restricted from being scraped, legally.
Just last year, the LinkedIn case: LinkedIn had a robots.txt, and the judge didn't give a fuck. Nor did he care what their terms of use said. Rather, it was hiQ's continued scraping of LinkedIn data even after LinkedIn's cease-and-desist letter to them that constituted access of data "without authorization."
>As to making a copy with intent to X, that’s what fair use is.
Yes, and?
>This distinction may seem crazy to you
It's not a crazy distinction, that's what I'd basically said previously, so perhaps we're talking past each other.
There’s a lot of precedent here showing scrapping isn’t guaranteed to be acceptable.
https://en.wikipedia.org/wiki/Facebook,_Inc._v._Power_Ventur....
$79,640.50 in compensatory damages + $39,796.73 discovery sanction
I could go on, but I am not sure what exactly you’re trying to argue here.
>I could go on, but I am not sure what exactly you’re trying to argue here.
That public is public. If you leave the door open in the real world, someone CAN enter your home. If you host your image on a public webpage, they can scrape it. robots.txt is not a security measure, nor is it a contract that magically gives the right to scrape where it wasn't given previously, it is a gentleman's agreement that you can ignore if you want to be a dick about it, and know about the robots.txt. Ethically, that's wrong, but it is how it is.
Not to mention, as I came to understand while reading during this discussion, LAION wasn't even crawling: they were using a public commoncrawl dump to gather their images. commoncrawl had crawled the author's site previously. They just took that data and got image links out of it.
1. they weren't selling a dataset
2. the artist didn't "disable" scraping in any meaningful way, legally
3. linking to the image is not illegal, and they're justified to respond with an invoice in Germany to recover legal fees for this dumb copyright complaint
4. it may fall under fair use to download images and train neural nets on them, it may not be. it always depends on the context and the specific case.
pretty simple stuff.
No trespassing signs have legal weight even without a fence.
Read up on Thomson Reuters v Ross Intelligence.