The end does not justify the means.
I just want artists affected to be paid, and actually paid. Not "paid" the way spotify artists are. if you're making a billion off of 100 artists' work, you better be making each affected artist a millionaire in royalties.
With AI, the situation is far from being that clear-cut. For one, the majority of datasets aren't comprised of capital-A Art with specific authors and attributions - a ton of it is just random information from around the internet, or just downright junk data. A photo I posted for free could conceivably end up in that dataset. Then, when a model is trained on that dataset and that model is used to make other outputs, are those outputs actually derivative with a direct link to some origin? If I generate a landscape using an AI, to whom exactly would I even owe money? Equally to all contributors in a dataset?
Random means there can be copyrighted content in it. It shouldn't be "random". You can't do such things when making your own piece of media, so I don't see why it's okay when it's a bot scraping massive pieces of the internet for use in some for-profit venture. Scraping was already this gray area in the internet and practices like these argue against such techniques.
>Then, when a model is trained on that dataset and that model is used to make other outputs, are those outputs actually derivative with a direct link to some origin?
That's the million dollar question (literally. Or more like 10/100 million). I imagine like existing copyright it depends on how far out you derive. Arguments above that suggest the ability to regurgitate the exact or very close to piece of data would certainly lean towards "yes". In which case it implies they have that piece of data in their database, instead of just being a "reference".
That feels too close hosting off of copyrighted material in my eyes, but I don't have the full picture of how and what these LLMs contain.
>If I generate a landscape using an AI, to whom exactly would I even owe money?
Like any media creator/producer would tell you: depends on thr license. For example, I Google "mountain" and my first result is this :
https://unsplash.com/photos/aerial-photography-of-mountain-r...
License says "free to use under the unsplash license", which is surprisingly lax. But there are two stipulations
Photos cannot be sold without significant modification.
Compiling photos from Unsplash to replicate a similar or competing service.
So we come back to the whole issue in the beginning with the first point. The 2nd point is a much more nebulous one to consider but not to be completely dismissed (maybe an AI service can be argued to generate a competitor).And that's one of the more generous results. Another result came from CNN but sourced from Shuttershock. Here's their license :
https://www.shutterstock.com/license
I'm not going to read nor summarize the license, but I'm sure CNN at the bare minimum needed to ask permission and likely negotiate to use it in their news post.
So yeah, an "ethical" training model will be doing this for every single image they use, if the aren't creating nor taking the pictures themselves.
>Equally to all contributors in a dataset?
Don't know. That would be for the courts to decide. If it's anything like streaming music and movies, you would be compensated proportionally to the amount of "usages" your work gets when generating pieces. Doesn't necessarily have to be equal, but it may surprisingly equalize out since algorithms are at the helm and not brands trying to stand out.
It's certainly be it's own rabbit hole to explore though. If we ever get that far.
[0] In the most egregious case, they found that just asking ChatGPT to repeat a word over and over again will make it spit out training set data verbatim. OpenAI's response was to stop any conversation with over a page of repeated words, so you can't replicate this.
which is fine in my books. You could equivalently produce the same works from searching thru the digits of pi.
The only enforcement that's needed is on the end user of the model - if they choose to produce the training set, they are violating copyright. The creator of the model _does not_ violate copyright merely by creating and distributing the weights of a neural network, as long as the space of potential output is vastly larger than the training set.
In this context, I don't see a principled difference between the model weights and really good compression. If I send you a gzipped copy of the latest bestseller book it's still copyright infringement. And it would still be infringement if I shipped it inside a software program that can _also_ reshuffle the words in a bajillion different ways, if there's a "copy" of the original work in there.
OpenAI and other large language model developers are claiming they have a machine that can write books, but they also fed it shittons of books, and they can't account for where all that text went. At best they can say "well, it doesn't produce exact, verbatim copies of the training set all the time".
in the same way you can brute force an MD5 Hash if you had a few centuries, I guess. I don't think "monkey's making Shakespeare" is a good metric for when to determine if a piece of art is unique enough to be scraped.
Tracing is looked down upon in the art community and this feels way too close to tracing if they just have these images in a database ready to reference. You can't store copyrighted movies nor music on such private databases (one of the few times I will ever utter the words "thank you DMCA"), I don't see why art pieces would be exempt.
I don't think fair use saves us either, at least not the model authors, because the whole selling point of these models is to replace artists. Yes, artists can use them as tools, but that is far less lucrative. The valuations and hype being thrown around specifically come from, among other things, being able to cut the creative class out of their own business. No court is going to look at that and say "ok, yeah, sure, it's perfectly fine for you to be using other people's work to train on".
An individual using an AI art generator may still wind up getting novel output that isn't obviously infringing to anything in the dataset. If that's the case then they probably haven't infringed copyright. Or at least, it'd be difficult to make a case around it.
Regardless, even if these cases did happen often enough, it's erroneous to assume that this is something universal (i.e. can be applied to all generative AI models) or intentional. Said models are vastly smaller than the sizes of their training datasets, so it's more or less impossible for all the data to be stored verbatim. Some aspects may appear to look like memorization if the piece of data reappears many times in the dataset - in those cases, the algorithm has a really strong incentive to recite that data. This reduces the effectiveness and is considered an artifact that needs to be corrected, not some underlying idea of AI.
The way these algorithms work is public knowledge, there's not really any black boxes in the hands of OpenAI that would be relevant here. Considering that, I'm really doubtful that these claims can be easily supported.
You appear to be refuting a slightly different point, though. When OpenAI was making Dall-E 2, they found that duplicates in the training set would incentivize memorization of specific images, as you said. This is more like finding a secret cheat code that would make the model regurgitate everything it had remembered, regardless of how "incentivized" it is to do so or even if it had been aligned to not do that.
My personal argument is this: the primary metric that the training process attempts to minimize is perplexity. This is how good the model is at guessing the training set data. Base models are specifically being designed to compress huge amounts of text and we just so happen to accidentally get a decent word calculator out of it. The alignment fine-tuning that happens later adjusts the model to prefer answering questions, but the underlying memorized data is still there.
>The way these algorithms work is public knowledge, there's not really any black boxes in the hands of OpenAI that would be relevant here.
Nope. Modern GPT is entirely a black box. OpenAI stopped publishing model weights the moment Elon Musk stopped writing the checks. Hell, they don't even publish the model architecture anymore. How GPT-4 works, even on a basic "this is how many transformer layers and attention heads we're using" basis, is a trade secret.
Even if you have model weights, the actual meaning of the learned model parameters has never been known; there is active research on figuring them out. One particular problem is polysemanticity. If you look inside a particular hidden layer, you won't see a single "dog" or "cat" neuron in its hidden layers. You'll have a 512-dimension concept bouillabaisse with "dog", "cat", "bird", "guinea pig", "kangaroo", and so on all floating around whatever shape made sense at training time (even if it implies absurdities like "desk is the opposite of loin cloth"). To untangle this, you have to train another AI to pick out monosemantic clusters of neurons that can then be inspected, and that requires extreme amounts of GPU resources.
Adobe has not shown how they train the text encoders in Firefly, or what images were used for the text-based conditioning (i.e. "text to image") part of their image generation model. They are almost certainly using CLIP or T5, which are trained on LAION2b, an image dataset with the very problems they are trying to address, C4 (a text dataset similarly encumbered) and similar.
I welcome anyone who works at Adobe to simply answer this question of how they trained the text encoders for text conditioning and put it to rest. There is absolutely nothing sensitive about the issue, unless it exposes them in a lie.
So no chance. I think it's a big fat lie. They'd have to have made some other scientific breakthrough, which they didn't.
Using information from https://openai.com/research/clip and https://github.com/mlfoundations/open_clip, it's possible to investigate the likelihood that using just their stock image dataset, can they make a working text encoder?
It's certainly not impossible, but it's impracticable. On 248m images (roughly the size of Adobe Stock), CLIP gets 37% on ImageNet, and on the 2000m from LAION, it performs 71-80%. And even with 2000m images, CLIP is substantially worse performing than the approach that Imagen uses for "text comprehension," which relies on essentially many billions more images and text tokens.