Getty Images v. Stability AI – Complaint
copyrightlately.com
copyrightlately.com
Getty announced in Oct they partnered with BRIA (?) to provide generative AI tools using their licensed images [1], and Shutterstock announced a partnership with OpenAI [2].
So it's clear these rights holders are OK with generative AI, as long as they continue to extract their pound of flesh. The language around "protecting artists" is horseshit - if you're a creative and you see Disney, Getty, etc getting behind your cause, you should look _very_ carefully around and make sure you're not the one being screwed.
1 - https://newsroom.gettyimages.com/en/getty-images/bria-partne... 2 - https://www.theverge.com/2022/10/25/23422359/shutterstock-ai...
> Rather than attempt to negotiate a license with Getty Images for the use of its content, and even though the terms of use of Getty Images' websites expressly prohibit unauthorized reproduction of content for commercial purposes such as those undertaken by Stability Al, Stability AI has copied at least 12 million copyrighted images from Getty Images' websites, along with associated text and metadata, in order to train its Stable Diffusion model.
Be compensated for the use of their content is maybe a more accurate and friendlier way to phrase this, but I guess not as edgy.
So yeah, not only is it very much Getty's content to distribute, but Stability AI is absolutely screwing over thousands of little guys by not paying royalties too...
It's difficult to find any source that would prove or disprove that. I can find references that they changed 35 million images to "royalty free" in 2014. The examples I can find of those are things like this: https://ichef.bbci.co.uk/news/976/mcs/media/images/73409000/...
Which appears to be a public domain image.
Or this $499 image from Getty: https://www.gettyimages.com/detail/news-photo/floyd-burrough...
Free from the Library of Congress: https://www.loc.gov/item/00651772/ (including the 50MB tiff that appears to be the highest resolution avaialable)
This is why Getty Images' claim is that Stable Diffusion is a "collage tool", while SD will most likely argue on the basis of the AI being capable of exhibiting the same level of creativity as a human.
It will still spit out stuff like watermarks and product photos unprompted, because its "learning" is still fundamentally mindless. It works strictly in the realm of pixels, it has no mechanism for understanding context.
So, not really.
You can't take a person and turn them into a cartoon if you're just pasting existing parts of images together. Stable diffusion understands what is a cartoon and it understands what is you (assuming you train it on your face).
It spits out watermarks because it understands that watermarks appear in images and it's going to try to reproduce the watermarks if you ask for something that tends to have a watermark in it.
With this in mind it will be even more interesting to see how this plays out.
Eg: https://www.gettyimages.com/detail/illustration/people-and-t...
This isn't a photo, but there are millions of pics as well.
Getty asserts a related point in the claim- that the work done to build extensive metadata for the content (which Stability apparently used for training their models) is a valuable service in and of itself.
Edit: just clicked through your link above and should have done so before. While the points above still stand (imho), the presentation on the Getty site borders on misrepresentation in the implication that they are copyright holders. Leaves a bad taste in the mouth.
My understanding is that in general, artists aren't necessarily against generative AI, but their complaint is (partly) around a complete lack of consent of being a part of these training models.
"Protecting artists" is not something they claimed as if they were opposed to AI usage at all. In their press release they said:
> It is Getty Images’ position that Stability AI unlawfully copied and processed millions of images protected by copyright and the associated metadata owned or represented by Getty Images absent a license to benefit Stability AI’s commercial interests and to the detriment of the content creators.
It's clear from this that the issue was not that Stability is an AI company, but that it's unlicensed. Getting an exclusive license to the images is specifically what they pay contributors for. Having those copyrights infringed by a competitor makes the content less valuable to Getty, and disincentivizes Getty from paying for new content in the future.
So yeah, it's plausible that this behaviour from Stability could harm content creators. Not because it's AI, but because its just run-of-the-mill unauthorized usage.
Me thinks you are an industry guy. You should read about all of the failed/lost cases that get swept under the rug.
I don't have a clue whether Stability AI will win this, it depends on their exact algorithm and how much they rely on source content, but if I draw something similar to the modern equivalent of the 'Mona Lisa', no, you don't have a right to it unless you can prove that in court. No, you can't copyright a painter painting a woman, either.
But now that you brought it up -- reproducing intact watermarks for Getty in multiple images as shown in the lawsuit feels like maybe your "slightly inspired" argument doesn't apply here.
But in this case you're not drawing anything. A little computer factory is. That same factory was built using images it didn't own.
Every human artist that ever lived(to my knowledge), heard, or saw someone else create a similar piece of art, from which they were inspired. If I create a song right now, how is that any different than an AI doing the same from being trained on copyrighted music. Certainly my song will be entirely made up of elements i've heard before, however large or small. An ML model is doing the same thing. There is nothing truly original in art. Artists are just filter and amplifiers of what they've heard, seen, and like. Your copyright does not permit you to restrict others from being inspired by your work, or using it for inspiration.
> If I create a song right now, how is that any different than an AI
It's vastly different in the fact that you are a human and the other is more or less a program built for business. There is only one version of yourself and your output where the AI has unlimited copies of itself and is only stopped due to lack of processing power and electricity.
I think even AI companies know there is a big difference and is the reason they use non-profit and researcher datasets instead of their own.
They own the images.
Basically it's an AI tool that takes a copyrighted photo, and produces an AI-produced photo that is "conceptually identical" but not actually identical.
That is, given a photo of "an asian man lying in a grass field surrounded by a semicircle of driftwood and crows", it would produce a photo that had all of the same concepts, but just slightly different execution on each of them.
The man's face is still asian but clearly a different person. The driftwood is still in a semicircle, but the individual pieces are all different. The crows are still there, but arranged slightly differently. The grass is still grass, but no blades are the same.
So it's essentially cloning the idea/concept/vibes of the target image, but none of the actual "implementation details".
Does anyone have any intuition of the legal outlook on this?
On one hand, nothing is stopping me from seeing the copyrighted photo and then recruiting a similar looking model, setting up a photoshoot in a field with some stuffed crows, etc. I could replicate what the AI is doing. It would be work, but I could do it. The AI is just automating this.
On the other hand, the actual stated intention of this tool is to get around copyright. Seems sketchy.
I don't think it is necessarily enough to have simply started with a copyrighted work.
Imagine if Stable Diffusion was made illegal. Someone accuses me of using this illegal tool for an image that doesn’t look like anyone else’s image as far as the court is concerned for copyright. I put the image on my website. If the image itself is not at all infringing, then what is the evidence that Stable Diffusion was used? Should the police be issued a warrant to search my private property for proof that I used Stable Diffusion without a shred of evidence?
I don’t know how it is implemented for the software in question.
Here's an interesting answer apropos to all this: https://opensource.stackexchange.com/questions/7250/could-i-...
There is no such thing as “clean room painting” and it should be really obvious why that is…
>On one hand, nothing is stopping me from seeing the copyrighted photo and then recruiting a similar looking model, setting up a photoshoot in a field with some stuffed crows, etc. I could replicate what the AI is doing. It would be work, but I could do it. The AI is just automating this.
If you trace art by hand that is still copyright infringement. If you paraphrase a passage from a book by hand that is still copyright infringement.
What if I see some fine art, I, a non-artist, make a super-low quality recreation of it with crayons, give that + a verbal description to a different professional artist who has not seen the original, and have them "upscale" my bad drawing into new fine art.
Their art would be conceptually very similar to the original. Same layout, same concept, same vibes, same style (if my verbal description was sufficiently good) but all the details would be different. Is this still infringement?
I think artists and copyright intermediaries would like to have "wildcard" copyright, "draw a flower once, all flowers belong to you now", and it would be very bad for creativity if they got their way.
I don't understand how this could work. Are there any examples out there you can cite?
If you have case law examples, it would be useful to cite them, but in general this not true. It can true when the paraphrasing is substantially similar to the original work. It would not be true, for example, if you paraphrase using all different words and a lot less of them. Copyright only protects the fixed tangible expression of the work, not the idea behind it.
It also protects against people making modifications to these works. When people parphrase something they typically do so by taking the original work, swapping words with synonyms, and shuffling the order.
If you described the original image (or made CLIP do it) and fed the description to txt2img generation, then you'd probably be fine.
I think being able to selectively optimize for "magic" random seeds in the diffusion algorithm will be kind of critically important here. Different seeds can produce very different images given the same prompt.
What if an optimizer can find a seed+prompt combos that are just as good as cloning as img2img?
If you instead don't want to clone an image, you can just extract the CLIP image embeddings from it and use them to condition a generative model like Dalle 2, Midjourney or Karlo (open source). The CLIP embeddings extract really well the semantic meaning of an image.
Something about going image -> prompt -> image feels like it "subverts" this somehow, even if the prompt is hyper-optimized to recreate the original image.
Obviously, this is just my feel/impression of it, the real test is how a jury feels about it.
The next few years will be really interesting in exposing that this is a really massive gray area.
That's only because you understand the algorithm in the first example.
At a jury trial most, if not all, of the jurors will find both methods equally opaque and, I believe, will treat them as equivalent.
The vectors are much more opaque because there's no straightforward human equivalent.
Describing to a human.
Describing it to stable Diffusion is very different. If you ask for a starry night, you get van Gogh's starry night, sometime with the original frame.
A prompt can easily trigger it to make a derivative work from something in its training set, and it often does.
Any competent expert will tell that to the jury and easily demonstrate it again and again.
Trying to explain how that "isn't really a copy" by explaining AI concepts isn't going to win the day, not when they can SEE the copy.
I can start with van Gogh's Starry Night, get the prompt "starry night, van Gogh" and get Starry Night back.
I'm using Starry Night as an example because Stable Diffusion consistently reproduces it, sometimes with the original frame, even with vaguely related prompts.
I'd say the jury will be making the right decision. Especially if the original image was part of the training set.
A tool that specifically advertises itself at getting around copyright is going to have a very hard time in court…
On the other hand, if I was to take an existing image had and then used that to transform it, it is a derivative work. If this was my image created with a camera rather than mspaint, its still a derivative work but I, as the creator of the original, can create derivative works of my own (I do it all the time).
The idea that I could use someone else's image and run it through img2img to make it "not a copy" misses the point of copyright and the exclusive right of the copyright holder.
https://www.copyright.gov/circs/circ14.pdf
> Only the owner of copyright in a work has the right to prepare, or to authorize someone else to create, an adaptation of that work. The owner of a copyright is generally the author or someone who has obtained the exclusive rights from the author. In any case where a copyrighted work is used without the permission of the copyright owner, copyright protection will not extend to any part of the work in which such material has been used unlawfully. The unauthorized adaptation of a work may constitute copyright infringement.
---
With starting from text, a description of the Mona Lisa is difficult to get generative art to create without significant effort to get the Mona Lisa (without saying "I want that image").
The question is then "is a description of the Mona Lisa a derivative work" and that is likely sufficiently far away from the original to be "no."
I think AI is overhyped in general, but as a tool to rapidly instantiate absurd hypotheticals it is really impressive. This is cool and good, IMO.
Okay, when you're advertising your product as a "get around copyright", you're going to so torpedo your credibility before the judge that there's no point trying to analyze the fair use factors--the judge is going to do whatever it takes to make them come out not in your favor.
If I take an image and then create a thumbnail of it would that be an infringing derivative work?
https://www.eff.org/deeplinks/2006/02/perfect-10-v-google-mo... and https://web.archive.org/web/20060813093554/https://www.eff.o...
Edit: on second thought, they might not even be derivative works (as a derivative work requires the same spark of creativity necessary for a copyrightable work), just outright reproduction. But the point that they are fair use still stands.
Note: the case you linked was appealed, and that ruling was reversed. https://www.eff.org/deeplinks/2007/05/p10-v-google-public-in...
> Fortunately, the Court wasn't buying it. It rejected Perfect 10's theory and found that until Perfect 10 gave Google actual knowledge of specific infringements (e.g. specific URLs for infringing images), Google had no duty to act and could not be liable. It also held that Google could not "supervise or control" the third-party websites linked to from its search results, something most people (except apparently Perfect 10) probably already knew. The rule provides strong guidelines for future development and avoids the kind of uncertainty that could chill start-ups trying to get the next great innovation off the ground.
> We conclude that the significantly transformative nature of Google's search engine, particularly in light of its public benefit, outweighs Google's superseding and commercial uses of the thumb- nails in this case. In reaching this conclusion, we note the importance of analyzing fair use flexibly in light of new circumstances. We are also mindful of the Supreme Court's direction that "the more transformative the new work, the less will be the significance of other factors, like commercialism, that may weigh against a finding of fair use." Campbell, 510 U.S. at 579.
> With respect to the second factor, "the nature of the copy- righted work," 17 U.S.C. § 107(2), our decision in Kelly is directly on point. There we held that the photographer's images were "creative in nature" and thus "closer to the core of intended copyright protection than are more fact-based works." However, because the photos appeared on the Internet before Arriba used thumbnail versions in its search engine results, this factor weighed only slightly in favor of the photographer.
> Here, the district court found that Perfect 10's images were creative but also previously pub- lished. The right of first publication is "the author's right to control the first public appearance of his expression." Because this right encompasses "the choices of when, where, and in what form first to publish a work," id., an author exercises and exhausts this one-time right by publishing the work in any medium. See, e.g., Batjac Prods. Inc. v. GoodTimes Home Video Corp., 160 F.3d 1223, 1235 (9th Cir. 1998) (noting, in the context of the common law right of first publication, that such a right "does not entail multiple first publication rights in every available medium"). Once Perfect 10 has exploited this commercially valuable right of first publication by putting its images on the Internet for paid subscribers, Perfect 10 is no longer entitled to the enhanced protection available for an un- published work. Accordingly the district court did not err in holding that this factor weighed only slightly in favor of Perfect 10.
...
> Having undertaken a case-specific analysis of all four factors, we now weigh these factors to- gether "in light of the purposes of copyright." In this case, Google has put Perfect 10's thumbnail images (along with millions of other thumbnail images) to a use fundamentally different than the use intended by Perfect 10. In doing so, Google has provided a significant benefit to the public. Weighing this significant transformative use against the unproven use of Google's thumbnails for cell phone downloads, and considering the other fair use factors, all in light of the purpose of copy- right, we conclude that Google's use of Perfect 10's thumbnails is a fair use. Because the district court here "found facts sufficient to evaluate each of the statutory factors . . . [we] need not remand for further factfinding." We conclude that Google is likely to succeed in proving its fair use defense and, accordingly, we vacate the preliminary injunction regarding Google's use of thumbnail images.
---
And so, is the model that Stability created significantly transformative, fundamentally different, and likely to provide significant benefit to the public?
> Google automatically makes low-rez thumbnails of all the images it indexes. The court concluded that Google's creation and display of these thumbnails from infringing websites did not fall within fair use.
> Having undertaken a case-specific analysis of all four factors, we now weigh these factors together "in light of the purposes of copyright." In this case, Google has put Perfect 10's thumbnail images (along with millions of other thumbnail images) to a use fundamentally different than the use intended by Perfect 10. In doing so, Google has provided a significant benefit to the public. Weighing this significant transformative use against the unproven use of Google's thumbnails for cell phone downloads, and considering the other fair use factors, all in light of the purpose of copy- right, we conclude that Google's use of Perfect 10's thumbnails is a fair use. Because the district court here "found facts sufficient to evaluate each of the statutory factors . . . [we] need not remand for further factfinding." We conclude that Google is likely to succeed in proving its fair use defense and, accordingly, we vacate the preliminary injunction regarding Google's use of thumbnail images.
I don't think it's that simple- as I understood the proposed conversion, there would be two phases. The first phase extracts an 'idea prompt' from a source expression (resulting in "an asian man lying in a grass field surrounded by a semicircle of driftwood and crows"). The second phase generates a new expression from this prompt alone.
As long as this intermediate 'idea prompt' is sufficiently devoid of 'expression' elements to withstand scrutiny (and the tool could even embed its idea prompt in metadata for auditing purposes), I would imagine the final output to be likewise considered a sufficiently transformative work compared to the original.
Edit: After gasp researching (which maybe we should all consider before commenting) you can patent one of 4 things, process, machine, article of manufacture, or composition of matter. Source: https://en.wikipedia.org/wiki/Method_(patent) so process is definitely something you can patent, but is not necessarily required.
Besides via collusion with the appeals court(s), this isn't going to happen.
However, the case was settled and the creator of the poster lied in court about his sources--so I'm not sure I'd draw too much about the poster inspired by photograph.
Weird Al gets permission for all of his music. Even if the parody defence were valid, song licensing is pretty cheap, and defending lawsuits against artists signed with major labels is expensive.
The problem here is that it does more than just clone the idea/concept/vibes, it really does tread into copying the implementation details. It matches lighting & composition, it matches subject and color, it can mimic the equipment used & props. People have done this manually, and been sued for it. Mostly it happens when an unknown artist steals the style of a specific well-known, best-selling artist. But now we’ve built a machine to near-copy anything in any style, with the intent of borrowing as much of the expression as legally possible, which seems like it probably can’t end well from a legal perspective. And because the technology for building these kind of machines is essentially public knowledge now, it’s hard to imagine this won’t be a problem from now on.
MIT Tech Review reports research with hundreds of similar results [1]. "The researchers, from Google, DeepMind, UC Berkeley, ETH Zürich, and Princeton, got their results by prompting Stable Diffusion and Google’s Imagen with captions for images, such as a person’s name, many times. Then they analyzed whether any of the images they generated matched original images in the model’s database. The group managed to extract over 100 replicas of images in the AI’s training set. "
[0] https://news.yahoo.com/researchers-prove-ai-art-generators-2...
[1] https://www.technologyreview.com/2023/02/03/1067786/ai-model...
Again, showing zero of anything resembling abstract understanding; merely a statistical correlation of blobs within each image with the text, e.g., finding the common set of pixels in each image that corresponds to "astronaut" and filtering out the rest.
Obviously, this is far more useful than color-key-deletion or whole-image search. But, it isn't intelligent abstraction.
What you are proposing is that you just "wash off" the expression from the idea and regenerate a new image from that idea. Great, except this isn't how AI art generators work. They aren't breaking down images into their core ideas, because those only exist in our human minds[1]. They're finding patterns of pixels that happen to match the text prompt well enough; and often times that includes the original image itself. Overfitting is a huge problem with conditional U-Net models and Google even released a paper detailing a way to find and extract memorized images out of an art generator.
So what will likely happen is that the art generator will just copy the image, or make one that's close enough that a judge would say that that it's a copy.
[0] If an expression is fundamentally wrapped up in an uncopyrightable idea and can't be expressed any other way, then it's also uncopyrightable. But if an expression is made up of uncopyrightable ideas, but separable from them, then you get a thin copyright on the arrangement of such.
[1] And, also, most humans are terrible at distinguishing idea and expression in the way that copyright law demands.
Which is a good thing, since having to read thousands of long spurious AI-generated copyright defences will quickly motivate lawyers to create laws against AI-generated legal defences.
That is somewhat analogous to how humans decide. Our brain runs over data and various potential outcomes and maybe reasoning and then decides. A bit later it generates a post-how explanation of said decision.I think we all have experienced ourselves and others having a totally wrong explanation about decisions but it is some how more satisfying than “a gut feel” or “because I said so”.
After that case there has been multiple theories on how to evade copyright law which all seem like they would equally fail at convincing a judge. One of my favorite is the method used by freenet, which takes a file and first encrypts it and then splits it into so small parts. Those parts are so small that multiple files will share identical parts with each other, so it is impossible to know for sure which file a person is downloading by just looking at the parts. In a different channel they also provide a recipe in how to reconstruct the file, and recipe by themselves are not enough as evidence to prove a download.
Sounds perfect until one would have to try convince a judge that no copying has occurred.
Getty Images is suing the creators of Stable Diffusion - https://news.ycombinator.com/item?id=34411187 - Jan 2023 (83 comments)
Others?
Maybe it is. A human artist rolling a die to determine color changes, what to include and exclude, etc isn't much different.
https://en.m.wikipedia.org/wiki/Monkey_selfie_copyright_disp...
https://www.lexology.com/library/detail.aspx?g=6c52581e-d82f...
Also NAL, but I'm cynical enough to believe that Getty's lawyers would avoid answering this question directly. And then wax lyrical about how their client should indeed receive a royalty for anyone attempting to use Getty's copyrighted works to learn the art.
Does it?
I mean, it probably does now, but did it say that at the time this training of Stability AI's model was going on? Did Getty have that foresight?
Most TOS boilerplate typicall prohibit commercial usage of their library without explicit license , Getty and every other company has that foresight if there is money to made they would want their cut is all that really need .
Let’s say you even just wanted to consume all images just sell an analysis of how many b/w images are there in their catalog it would still breach of their terms unless their TOS allowed you to do so the novel copyright question may simply not even matter in this particular case.
So in some ways, you can argue that Stability is also directly redistributing the original images (albeit in a compressed format).
> Making matters worse, Stability AI has caused the Stable Diffusion model to incorporate a modified version of the Getty Images’ watermark to bizarre or grotesque synthetic imagery that tarnishes Getty Images’ hard-earned reputation, such as the image below
(see page 18 for an example)
Getty is going to win something. There is clearly a problem with the model. The outputs are often not novel enough to make them indistinguishable from the training data.
Both were built by scanning copyright materials, and in Authors Guild the Southern District of NY found that such scanning does not constitute a copyright violation.
I don't think so and the complaint isn't just about 'trademarks' either.
OpenAI was able to get explicit permission [0] from Shutterstock to train and on their images for DALLE-2. Stable Diffusion did not and is commercializing the use the model with Dreamstudio as a SaaS offering which the model has found to be outputting images with Getty's watermark [1] without their permission. That doesn't seem to be 'fair use' nor is it transformative either given the watermark is clearly visible in the generated examples here: [1]
This is going to end with a settlement and Stable Diffusion licensing deal with Getty over the images; just like with OpenAI did for DALL-E 2 with Shutterstock. Neither Shutterstock or Getty are against Generative AI either even as shown in this deal with Getty recently [2]
[0] https://www.prnewswire.com/news-releases/shutterstock-partne...
[1] https://www.theverge.com/2023/1/17/23558516/ai-art-copyright...
[2] https://newsroom.gettyimages.com/en/getty-images/bria-partne...
Sure, they also feed on each others work etc but in the core of all these copyright, piracy, patent and similar discussions is how these people are supposed to be compensated.
Working in the software company in the day and preaching open source, anti copyright anti patents open access free for all in the night works for the software people but people in the creative industries are really struggling to get paid for their work.
The genie isn't going back in the bottle, the tech will be able to produce derivative work over the work of other people and I'm not looking forward for the greater number of struggling artists.
You don't need tech (or at least computers).
To the degree that a portrait photographer, say, has a distinctive lighting and posing style, that can absolutely be copied. And there are many examples in art of art techniques that were widely copied.
It's like building your security on hard to brute force secrets in tech and suddenly someone makes a machine that instantly brute forces any secret. Its a similar kind of disaster with the difference that human being can't just switch doing something else and the value they added to the society is not compensated.
"Open source" is copyright - it's not anti-copyright. It uses copyright to grant a license to use under certain conditions, and sometimes with obligations. You might keep it proprietary, you might use GPL to require that the software stays open virally, or you might use a more permissive BSD-style license. The important part here is that as the creator, you choose how you want your work to by copyrighted.
[0]: best quickest link i could find that contains the "consent, credit, compensation" https://mindmatters.ai/2023/01/three-artists-launch-lawsuit-...
Trainers should require consent from artists to train their model on an artist's work. A part of obtaining that consent could be some form of compensation and ideally credit when generating the images. I don't believe many artists are necessarily concerned about people copying their style. From what I've seen is they just don't want their artwork and their style sucked into the AI-borg to be reproduced en masse.
I do not think scraping images and using them to train models is fair use. I believe AI labs should obtain consent.
The AI will launder the content just like GitHub's CoPilot, and attribution will be impossible on the other end. Since Creative Commons licenses are not PD, and often do require attribution (CC-BY) or they prohibit commercial usage (CC-NC) or they require that derivatives must be licensed the same way (CC-SA) or even prohibit derivatives outright (CC-ND) all of those requirements are going to be stomped into dust by generative AI.
And those licensors won't be big enough to sue anyone.
But I would argue that Stable Diffusion with the open-sourcing of their model weights, and use of the LAION dataset which is released under CC-BY 4.0, would likely meet both the letter and intent of the license. https://wiki.creativecommons.org/wiki/CC_Attribution-ShareAl...
Like I said before, I believe there's a strong argument that Stability AI / LAION's use of CC-{BY/SA/ND} is likely allowed under the terms of all CC licenses due to the works being shared without alteration and with attribution, released under CC-BY-SA 4.0 (LAION) and the Stable Diffusion model being released under a permissive license (CreativeML Open RAIL-M).
The real question is if the images generated by the models need to provide attribution to every single weight involved in generating that image. That's a lot more complicated and unclear, but quickly gets into questions like "Should artistic style be copyrightable?" and "What amount of source material is required to constitute a copyrighted work?". But as of right now, I don't see how any of this is violating the letter or intent of CC 4.0
If you haven't registered your copyright, you can sue for actual damages which includes the money you lost (including potential money) and maybe the money gained by the person using the work.
You will need to prove that the work generated by stable diffusion is an infringement of your work and that stable diffusion is liable (this will be challenging). You could also try suing the person using the work for profit (it will need to be for profit because otherwise all three points of "what you can sue for" is $0).
Remember to register the images that you create with the appropriate copyright office if you wish to be able to sue and have some teeth (and have a better than zero chance of collecting lawyers fees from the other party)
This can be done as a "group registration" to do bulk and doesn't need to be done one at a time - https://www.copyright.gov/registration/photographs/ and https://www.copyright.gov/registration/visual-arts/
I put my music online under a CC license because I want people to be able to play it freely, use it in videos, remix it, cover it, include it in a compilation. I'd like for people to be able to do anything with my music except claim to be the original creator.
They're selling knitting needles in the era of robotic manufacturing.
I hope every single one of these lawsuits falls flat on its face. Other countries will happily overlook Getty copyright to get the leg up on AI.
AI is not reusing copyrighted material. It's learning from it in the same way humans do. You can even fine tune away from the base training set and wash any experience of it away.
Besides, if Getty wins, it merely insures that the large incumbents with massive pocketbooks to pay off Getty et. al. win. It'll keep AI out of the hands of the rest of us.
Wedding photographers make a lot of money. Sports, events, local artists... There's plenty of money in photography.
All this science fiction is right; there isn't room on the planet for both humans and AIs. It sounds depressing that a bunch of GPUs are going to kill us all off, but it was coming anyway. The sun becomes a red giant and consumes the Earth. All protons in the Universe decay in 10^17 years, ending the existence of matter. The trajectory is clear even if the means aren't; humanity can't last forever.
If I sound depressed, I'm not really. People just use the headlines to guide their view on what The End looks like. Read a few articles about chatbots, and it's AIs taking all our jobs. Watch a few movies about asteroids, we go out like the dinosaurs. Hear "Russia invades Ukraine" and it's a nuclear holocaust. Read a few particle physics papers, and it's proton decay. You can't worry too much about it. Enjoy your time while you have it!
On the other hand humans are self replicators and only need a bit of biomass for sustenance, biomass that grows by itself, too. No factory, no supply chain, we got everything we need to make more of us.
If you consider the risk of EMP, an AI needs humans to restart it, or some way to survive electronic attacks.
for now
> If you consider the risk of EMP
if you consider the risk of a bioweapon ...
This assumes 1. automation is free, 2. humans cost too much, so any company would ditch their humans for AI. But in reality AI costs money, AI is better with people than without, and people can generate profits surpassing the cost of their wages. Why would a company prefer to reduce costs to increasing profits? When everyone has AI, humans are the differentiating factor.
What will AI learn from if Getty no longer exists?
This is like crying over Rolodex.
And let's not forget how awful Getty has been throughout its existence. They've frequently sued people for things they didn't even own the copyright to.
Novelty of subject will likely get covered by partially unwitting data gatherers (eg google photos)
They're teachers in the age of students.
> We don't need to prop up the old at the expense of the new.
We don't need to pay teachers when we can just profit by charging students tuition.
This is a very good quote, but unfortunately I fail to grasp it. Care to elaborate?
>Without new images to learn from
Even in the event that all cameras are destroyed (somehow), people will use generative models to describe their experience in a new era, and then this new knowledge will be used by new models and the cycle will repeat itself.
Yeah...no. AI is doing nothing but reusing material. It generates the most likely image/text/code in its training set to be found following/around/correlating with the prompt. It literally has nothing outside it's training set to reproduce. And when it reproduces the Getty watermark, that's pretty obvious example of reusing copyrighted material.
>>It's learning from it in the same way humans do.
Not even close. These "AI" architectures may be sufficiently effective to produce useful output, but they are nothing like human intelligence. Not only is their architecture vastly different and making no attempt to reproduce/reimplement the neuron/synapse/neurotransmitter and sensory/brainstem/midbrain/cerebrum micro- and macro-architectures underlying human learning, the output both in the good and the errors is nothing resembling human learning. (source: just off-the-top-of-my-head recollections from neuroscience minor in college)
Yikes.
This is simply false. It's not a search engine that outputs the training item closest to the prompt.
In reality, it is "learning" (in some sense) how to correlate text to images, and then generating brand new images in response to input text. If this is legal for humans to do, then it's probably legal for machines to do the same thing.
>>It's not a search engine that outputs the training item closest to the prompt.
Correct, it is not outputting the training ITEM, it is outputting finer-grained slices of many items, more of a mash-up of the training items.
Of course it is not taking an entire specific image the closely matches the search term, it is taking averages of component images of "astronaut riding a horse over the moon in style of Rembrandt".
That image won't exist in the training set, but astronauts, horses, and Rembrandt-style coloring and shading do exist, and it is assembling those from averages of the components found it's training set, not from some abstract imagination or understanding.
The fact that the astronaut suit may not be the exact same as any of it's training images is the same as if I averaged 100 faces in photoshop, not because there is some kind of "learning" or "understanding". Ability to do useful statistical mashups is NOT the same as "learning".
This can be shown in a different "AI" engine'd failure to solve a child's puzzle. ChatGPT, when presented with: "Mike's mom had four kids, three are named Lucia, Drake, and Kelly, what is the fourth kid's name?". It said there is insufficient info, and doubled down when told that the answer is in the question.
>>how to correlate text to images
yes, as I pointed out, "correlating with the prompt." I didn't say it correlated an entire image, but I also failed to specify that it was correlating components.
>> If this is legal for humans to do, then it's probably legal for machines to do the same thing.
This [0] is I'm quite sure, not legal. Asked for an image of a person named "Ann Graham Lotz", it returned the image in the training set, slightly degraded.
First, that is literally the search engine functionality you were deriding.
Second, if you asked a human artist to produce the same image, without infringing copyright, they would produce something likely recognizable as the person, but obviously not resembling the training photo. It doesn't matter if they are a portrait painter, sketch artist, Photoshop jockey, or Picasso-like impressionist.
So, no, this does not represent learning in any conceptual, creative, or human-like sense.
It does represent mashing-up averages of inputs of various components. Feed in enough "astronaut" photos, and it'll be able to select out the humans in the spacesuit as the response to that prompt. Same for "horse", "moon", "riding", and "Rembrandt". and it can mash them together into something useful with good prompts.
But give it something very specific, like a person's name, and you get basically a search-engine result, because it doesn't have enough input data variety to abstract out the person 'object' from the background.
[0] https://techxplore.com/news/2023-02-ai-based-image-generatio...
> it is assembling those from averages of the components found it's training set, not from some abstract imagination or understanding
how exactly are you so certain that the human brain handles abstract concepts any differently? please note that I'm not claiming that I myself know, but rather that you almost certainly do not know and thus are presenting an invalid argument
what is human imagination anyway?
> assembling those from averages of the components found it's training set
> slices of many items, more of a mash-up of the training items
> But give it something very specific ... it doesn't have enough input data variety to abstract out the person 'object' from the background
so is it abstracting or not? where's the line between that and a mere statistical mashup?
Good question. At the very least, we have a far deeper understanding of physical reality. Humans would not unintentionally (e.g., for effect) produce images of people with three ears, or of a bikini-clad girl seated on a boat with her head and torso facing us, and also her butt somehow facing us and thighs/knees away... yet I've seen both of these in the last week (sorry, couldn't find the reference, it was a hilarious image, looked great for 2sec until you saw it)
I admit that it is possible (tho I think unlikely) that this is a difference in quantity, not in kind.
One reason to doubt this is that Stable Diffusion was trained on 2.3 billion images. This is a vastly larger library than any human has seen in their lifetime (considering that viewing 2.3 billion images at one per second would take 72.8 years). Yet even if you count every second of eyesight as 'training', children under 1/10 of that age, who have seen only 10% of those images would not make the same kinds of mistakes.
Plus, the neuron/synapse/neurotransmitter and brainstem/midbrain/cerebellum micro & macro-architectures are vastly different than the computer training models. So, I think we can be confident that something different is happening.
>>so is it abstracting or not? where's the line between that and a mere statistical mashup?
Good question. There is definitely something we might call, or that might resemble abstraction. It's definitely able to associate the cutout images of an astronaut in a spacesuit from the backgrounds. It can evidently assemble those from different angles.
But it certainly does not have the abstraction to understand even the correct relationship between the parts of a human. E.g., it seems to keep astronauts' parts in the right relationship, but not bikini-clad-girls' parts (because of the variety of positions in the dataset?). There's no understanding of kinesiology, anatomy, or anything else that an actual artist would have.
Could this be trained in? I expect so, but I think it would require multiple engines, not merely six orders of magnitude more training of the same type. Even if 10^6X more training eliminated these error types and even performed better than humans, I'm not sure it would be the same, just different and useful.
I'd want to see evidence that it was not merely cut-pasting components of images in useful ways, but generating it from an understanding of the sub-sub components: "the thigh bone connects to the hip bone, the hip can rotate this far but not that far, the center of mass is supported...+++" as an artist builds up their images. Good artists study anatomy. These "AI"s haven't a clue that it exists.
>>to me the "search engine" case where it reproduces a specific training image seems like a failure mode that's distinct from normal operation
Au contraire, it seems that this merely exposes the normal operation. Insufficient images of that person prevented it from abstracting the person components from the background, so it just returned the whole thing. IDK whether it would take a dozen, hundred, or thousand more images of the same person, to work properly. But, if they all had some object in the background (e.g., a lamp) that was the same, the "AI" would include it in their abstraction.
(but I could be wrong).
yes my point was that this total failure to abstract (or slice or average or whatever it is that it usually seems to do) appears to me to be neither the intended nor typical mode of operation
> children under 1/10 of that age, who have seen only 10% of those images would not make the same kinds of mistakes
but then children aren't being fed a stream of unrelated images. they're receiving a wide array of real time sensory input from an environment they're actively operating in
consider your examples of the lack of higher level understanding about how the parts of a human "fit together". what practical experience do these models have that could actually convey such an understanding? deriving a proper understanding of mechanics in 3D from one million independent 2D still frames of human hands performing various tasks seems like it should be extremely difficult at best
> Could this be trained in? I expect so, but I think it would require multiple engines
I think it requires a different sort of training algorithm entirely. work such as https://arxiv.org/abs/1803.10122 suggests to me that there might be little difference between the human ability to abstract and lossy compression. at the same time work such as https://arxiv.org/abs/2205.11502 makes it apparent that in many cases this sort of generalization simply does not happen the way we'd like
> the neuron/synapse/neurotransmitter and brainstem/midbrain/cerebellum micro & macro-architectures are vastly different than the computer training models. So, I think we can be confident that something different is happening
something being architected differently doesn't necessarily mean that the higher level functionality is any different
moreover, in purely functional terms how do you propose to distinguish something that's different from something that's incomplete? ie a smaller piece of a larger whole? if someone constructs for example a passable digital model of the visual cortex of the mouse or human or other animal that's still only a single small piece of the whole
so who is and how are we to say that we either have or haven't achieved a meaningful form of abstraction versus merely averaging bits of the training set together? at this point I'm not actually clear where the line between those two things even lies
Yup, certainly not intended, although I see it as the typical response on the edges of the data set; objects with too few varied representations will always fail in this way. Seems square-cubish as there will always be a volume of solid training data and a surface of partial data, so maybe not severe.
>> deriving a proper understanding of mechanics in 3D from one million independent 2D ...extremely difficult at best
Yup. This is definitely part of how it is different. Doing the full training set with stereographs would likely improve it, but it'd improve it even more to have the same images manipulated by robots and the feedback integrated. Considering the 3.5 billion parameters of DALL-E, 4.6B for Imagen and 890MM for Stable Diffusion, how many params would be needed to integrate stereo-vision and robotic feedback? 3.5billion squared or cubed? Would that be enough just scaled up, or do we need to qualitatively change the structure?
>>I think it requires a different sort of training algorithm entirely.
Agree 100%. I think these engines are a part of the solution, but not the whole. I expect we'll need multiple different kinds of training models, and then the methods to integrate them and correlate their 'knowledge'. E.g., figuring out how one part of a moderately complex object (e.g. a human) hides another part in certain positions (e.g., hand behind back) is trivial for a 3D modelling system, but even the massive 2d ones often get it wrong.
>>being architected differently doesn't necessarily mean that the higher level functionality is any different Definitely true. Parallel evolution, elec vs ICE powered cars, etc. The question is when we've achieved the same level of functionality.
>>how do you propose to distinguish something that's different from something that's incomplete?...achieved a meaningful form of abstraction versus merely averaging bits of the training set together? at this point I'm not actually clear where the line between those two things even lies
YES, excellent question. Especially since these models don't do much explaining of their inner workings. Humans also haven't fully figured out our inner workings either.
It's looking right now like different AI will arrive faster than biomimicry-based AI, partly because we still don't know the bio at a deep enough level. IDK if it'll stay this way.
I remember discussions a long time ago with a scientist who worked on AI for early Mars missions, and how they'd move their machines. He was describing the algos for tracking the world, their machine, and adjusting motion, with the team assuming that they were re-creating the way humans do it. From my experience as an international level athlete and a neuroscience minor in college (inspired by my sport experiences), I could tell that his methods were nothing like how biological systems work. Seeing Google's self-driving car drive around a racetrack was truly impressive, but from my sportscar-racing training 7 experience, I could instantly tell that it was accomplishing the task nothing like any human would, although it was achieving competent levels of performance (in a limited setting).
How do draw the line? It may come down to the kinds of clever tests built by childhood and animal behaviorists to study animals who can't self-report on their state or if they actually figure out something or not.
That said, I don't think it's impossible for an AI to end up exceeding our capabilities by using different methods. Kind of like Paul Bunyan vs the chainsaw.
(BTW, thanks for the lively discussion; it's a pleasure to be pushed to define my thoughts better, and I've learned; happy to keep it going)
Only because some people named their field "machine learning" and called it "learning".
It has no relation to human learning.
If your child accidently confuses a giraffe with some other animal you correct then, you don't add the picture to their training set and show them again thousands of pictures of giraffes hoping that their success rate improves.
If you ask Stable Diffusion for a starry night, you get van Gogh's starry night.
Getty is slightly more than just a website that posts low resolution, low quality pictures with a fat watermark on it.
Has nothing to do with robots or AI...
Stable Attribution - https://news.ycombinator.com/item?id=34670136
Of course copyright is an abstract legal tool, so mo argumemt is worth anything until it's codified into law/precedent.
This changes nothing related to copyright. Humans use tools all the time, and those tools aren't humans either. A camera isn't a human, for example, and yet it does a large amount of the work of making the photograph.
In regards to copyright, no matter what tools you use to create art, what matters is the output. Drawn by a Human, created via a non-human such as either a camera or an AI, it is all the same.
By the same vein of thought I could claim that Google is a tool and by typing in different keywords I can make it "output" almost any image. Yet me nor Google could never assert copyright over these images simply because I can type keywords to cause image search to output them, no matter how creative or original keywords I use.
Good question. The difference is regarding the output not the input.
Yes, if you output art, that looks exactly like some other art, then this could be a copyright violation. But that has nothing to do with a computer. Regardless, of if you copy someone else art, by right clicking it, AI generating it, or if a human entirely using their own paint and pencils, completely recreates it, this is all the same.
The computer has nothing to do with it. Neither does the input.
> no matter how creative or original keywords I use.
Exactly. Regardless of the input, the input doesn't matter. That is why AI art is legal. Because the complaints are not about the output, but instead about the input.
Yes, if you input other people's art, but the output is transformative, then that is legal.
Same for if a human does it, or a computer, or anything else. Human input, vs computer input is the same thing, whereas the illegal stuff, is based if the output is infringing, regardless if it is made by a computer or a human.
> By the same vein of thought I could claim that Google is a tool
Yes, it is a tool. And just like any other tool, if it is a computer, or a human, it is the output that is judged. Same as if a human took a picture, or hand copied someone else's art.
A human hand copying art, is just as legal or illegal, as if a computer does it.
> me nor Google could never assert copyright over these images
This is an unrelated topic. This is if the generation of new images, is copyrighted or not. That is different from if you are infringing on someone else's copyright.
I was only talking about "if you are infringing on someone else" topic, not about the generation of new copyright, which yes, requires human input, as according to the "monkey selfie copyright" case.
If the outputs are not original, AFAIK they then must be derivative. Stable Diffusion could claim fair use exemption. But fair use too is just meant to protect creativity, again a manifestly manual activity.
I don't know which way I lean, but I sure know the courts will soon have to make some very interesting rulings that will have monumental importance.
Maybe A.I. generated "art" is an entirely new class of work and lawmakers just need to rethink copyright for them.
Building a collection of stock images is the predecessor to generating images from text.
Also variations have to be extreme in order to not violate copyrights, and even then may still violate. Otherwise youtube would have no issue with me uploading all of Shrek as long as I mirrored the video and pitched the audio up by 3%.
Thats not illegal.
For example, right now, for whatever job you are doing, you have probably looked at copyrighted works, that you don't own.
You have copied those copyrighted works, because in order for you to view the image that you don't own, you had to download it to your computer.
So, you have thus used copyrighted works, for commercial purposes, if you have ever looked at copyrighted works on your computer, for a professional purpose.
What is infringement, is not using copyrighted works for a professional purpose. Instead it is distributing copyrighted works to other people that is infringement.
It's there.
It's public
Anyone can rent compute and as AI advances it's only going to get easier to fine tune models.
There's going to be a world where you can just point a program, one on your computer at a website with a bunch of images on it, wait 10 minutes, and I have it punching out stable diffusion style images.
What are the courts going to do to stop it? Sneak into your home and record everything you do? How are you going to prove an AI was trained on certain images? How are you going to prove an AI generated an image?
The cat is out of the bag, even if all the courts have decide that these images are copyright and can't be used, they're going to continue to be used by people, all over the place, with absolutely nothing being able to stop it.
The era of needing an artist to produce a novel image for a purpose is over. It will never come back, Even with the full support of the law trying to keep it around.
Just because it's easy to speed on a road and other people are speeding and they aren't charged doesn't mean you can't be held liable for speeding if you're caught. Likewise, copying proprietary source code from another project into your own commercial software can seem innocuous until you get hit with a lawsuit. You "prove" it the same way you prove any other thing before a court. The jury/judge doesn't need to be absolutely certain that something did or did not happen, just that the plaintiffs prove it beyond the standard of proof.
So yes: we may as well enter a point that computers are virtually indistinguishable from humans in generating "novel" things, in that case the existing laws would need to change. In the meantime, I don't view AI models as a trump card around copyright/IP. If everyone else is following the rules of the road but you decide to stick it and do your own thing, don't expect to ram your way through without consequences.
- https://petapixel.com/2016/11/22/1-billion-getty-images-laws...
- https://www.techdirt.com/2019/04/01/getty-images-sued-yet-ag...
- https://en.wikipedia.org/wiki/Getty_Images#Copyright_enforce...
Selling public domain things is entirely legal. Claiming they own those images and suing others for using them is not - they're public domain.
The article you linked makes the point that at least one civil suit did give damages for dealing with a bogus DMCA takedown, but that it was unrelated to the perjury provision.
The outcome will be, Stable Diffusion settling and licensing the images from Getty Images. If OpenAI was able to do it with Shutter-stock, so can Stable Diffusion.
1. The requested licensing fee approaches infinity
2. Getty Images simply refuses to license images to anyone who will use AI to create derivative works
You can use public domain images for profit. It's not surprising this was thrown out.
> It's comical that now it's exactly the same allegation, but with the sides inverted, now it's Getty trying to sue an AI company for using their public domain images.
Where does it say they're suing over the public domain images in their collection? Their collection is not entirely public domain images. Their suit claims for the copyright works by staff photographers, third parties that have assigned copyright, and images licensed to them by contributing photographers. In addition, they're claiming for the titles and captions which they created and are themselves copyrighted.
It's not the "exact same allegation", and there's really no relation between the facts of the cases here.
Getty’s strategy at the time appeared to be to meet any infringement accusations with a massive legal response. Any individual photographer could generally not afford to respond.
My impression was that they were not super concerned about infringing on other’s work. But they will sue the pants off anyone who they perceived to be violating their copyrights.
But in recent years there are legal firms dedicated to pursuing deep pocket infringement cases on contingency. This has changed the legal calculus for large companies who were not careful with copyright.
I'm surprised at this in your case because typically copyright infringement cases are massively weighted in favor of the copyright owner (defined damages AND reimbursement of legal fees). I know this because I'm currently a defendant in a copyright lawsuit.
Not sure about their reputation but they have a point with the bizarre/grotesque thing.
It might be much more complicated than it appears in the surface! For example, look up Richard Prince!
https://www.cnn.com/2015/05/27/living/richard-prince-instagr...
The earliest I heard of it was the "parsing html with regular expressions" classic which I highly recommend if you haven't seen it before: https://stackoverflow.com/a/1732454/4012132
𝐼 𝕒𝕝𝕤𝕠 𝓁𝒾𝓀𝑒 ㄒ卄丨丂 one: https://lingojam.com/VaporwaveTextGenerator
It's all just clever use of unicode.