fyi, for Adobe Firefly, we are training or models. From the FAQ:
"What was the training data for Firefly?
Firefly was trained on Adobe Stock images, openly licensed content and public domain content, where copyright has expired."
(I work for Adobe)
fyi, for Adobe Firefly, we are training or models. From the FAQ:
"What was the training data for Firefly?
Firefly was trained on Adobe Stock images, openly licensed content and public domain content, where copyright has expired."
(I work for Adobe)
The contributor agreement linked from here[1] is this: [2]
"You grant us a non-exclusive, worldwide, perpetual, fully-paid, and royalty-free license to use, reproduce, publicly display, publicly perform, distribute, index, translate, and modify the Work for the purposes of operating the Website; presenting, distributing, marketing, promoting, and licensing the Work to users; developing new features and services; archiving the Work; and protecting the Work. "
I guess this would fall under the "developing new features and services".
What is funny is that "we may compensate you at our discretion as described in section 5 (Payment) below". :) I like when I may be compensated :)
And in section 5 they say: "We will pay you as described in the pricing and payment details at [...] for any sales of licenses to Work, less any cancellations, returns, and refunds."
So yeah. Sucks to the artist who signed this. They can use your work to develop new features and services, and they do not have to pay you for that at all, since it is not a sale of a license.
1: https://helpx.adobe.com/stock/contributor/help/submission-gu...
2: https://wwwimages2.adobe.com/content/dam/cc/en/legal/service...
And in this case, to develop new features and services that specifically undercuts your existing business, viz. selling stock photos for money. Sucks to the artists, indeed.
From the popular Getty Images, used by media worldwide: > When a customer uses one of your files, it may appear in advertising, marketing, apps, websites, social media, TV and film, presentations, newspapers, magazines, books, or product packaging, among many other uses.
The vast majority of copyrighted works were conceived and negotiated under conditions where ML reproduction capabilities didn't exist and nobody knew what related value they were negotiating away or buying.
And on top of that, it will become a spotify where each creator gets a sum total of $0.00000000001 per AI their media item was trained on and maybe a few dollars a month, while paying a greater tax to apple-sony-disney whenever their AI style recognizers charge you a royalty bill for using whatever bullshit styles it notices in your media items.
Copyright should stay in it's 'exact duplication' box, lest we release an even worse intellectual property monster on the world.
This is a cartoon conception of both existing copyright and the proposal I described.
Copyright applies to works, not "ideas" (William Gibson doesn't have a copyright claim on the "idea" of a virtual reality but he does on _Neuromancer_ as a work). It gives creators incentives and a stake in their work by allowing for some degree of control in how works are used and especially how they're distributed.
What I'm proposing isn't that different, and it's simple:
People should be able to decide if their work is put into a training set, and what they get in return for it.
That's it. No "flavor copyrights." No ownership of ideas. Just a claim on the creator's work and how it's used/ distributed. Specifically updated to cover a case that didn't exist 10 years ago, a case whose explicit intention is reproduction of not just one narrow aspect of the work but a combinatorially large number of aspects of the work (there's no other reason for adding it to the training set, that's the nature of these models).
> it will become a spotify where each creator gets a sum total of $0.00000000001 per AI their media item was trained on
"the artists will get such a small payout so we should make sure they get nothing" is quite the take.
I think "artists should be able to negotiate what consideration they want in return for their work for a training set" is a better one. It's certainly better than "people collecting training data should be able to take anything they can get their digital hands on without any consideration other than possession." Maybe some artists would take a nanocent per use. Maybe they wouldn't. That should be up to them.
Part of the reason why services like Spotify are so terrible economically is the way digital rights were assumed by labels and streamers often as extensions of older pre-digital conceptions without much in the way of negotiation by artists. The current moment is a chance to do better across an equally large (if not larger) technological shift.
For example, even a short sample used in a song usually has to be licensed. Cover versions of songs may qualify for a compulsory license with a set royalty payment scale.
However some reuse (such as transformative use, parodies, or use of snippets for various purposes, especially non-commercial purposes) may be considered fair use. AI companies could very reasonably argue that use of images for training AI models is transformative and qualifies as fair use, that no components of the original images are reused in AI-generated images, and that AI-generated images are no more infringing than human-generated images which show influences from other artists.
Absent additional law, I expect the legal system will have to sort out whether AI-generated images infringe the copyright of their training images, and if so what sort of licensing would be appropriate for AI-generated (or other software/machine-generated) images based on training data from images that are under copyright.
Though this sounds extreme, enforcing the alternative would break any last remnant of human privacy. It would kill the independent operation of computing as we know it and severely cripple AI/ML research when we need it most: human alignment.
It is possible that a catastrophic event occurs and halts the supply chain of advanced semiconductors in the near future, in which case the debate can be postponed indefinitely.
Yep. They could. And people do.
But this particular use is entirely new, though. Old conceptions of fair use can't and shouldn't cover it. "Transformative"? Sure, but there's a difference between transforming a work once and laying hold of it in automation to be transformed on request millions or billions of times in degrees ranging from simple to convoluted. It's hard to argue that the right to this exists in fair use when awareness of the possibility didn't exist when fair use conceptions were constructed.
The output isn't the issue so much as the consent & consideration for use of the input. We'll need case law or statutory law that understands this. It should be possible, but it should be possible on an opt-in basis. When you're training a model, you should know.
> AI-generated images are no more infringing than human-generated images which show influences from other artists.
If we end up with a policy differentiating the two and privileging humans and works created by them, I'm fine with that.
The training is based on the licensing agreement for Adobe contributors for Adobe Stock.
(I work for Adobe)
But the naive approach of having a table of how much each individual training item influenced every weight in the model seems impossibly big. For DALL-E 2's 6.5B parameters and 650m training items, that's 4.2 quadrillion associations. And then you have to figure out which weights contributed the most to an output.
I would love to see any research or even just thinking that anyone's done on this topic. It seems like it will be important in the future, but it also seems like a crazy difficult scale problem as models get bigger.
> How can we explain the predictions of a black-box model? In this paper, we use influence functions -- a classic technique from robust statistics -- to trace a model's prediction through the learning algorithm and back to its training data, thereby identifying training points most responsible for a given prediction.
If I include the tag "floor", do I get some (tiny) percentage of every image that uses "floor" in the prompt, even if the bits from my image did not end up affecting model weights much at all in training?
Worse, for tags like "dramatic lighting", it's likely that the important source images will depend on the other words in the prompt; "sunset, dramatic lighting" will probably not use the rely on the same weights or source images as "theater interior, dramatic lighting".
And then you get the perverse incentives to tag every image with every possible tag :)
I'd love to be convinced otherwise, but I'm not seeing prompt-to-tag association working.
Why do you have to, though? What do you hope to trace back to exactly?
(Also, it's entirely possible that eg a model could generate images resembling your work without "seeing" any of your work and only reading a museum website describing some of it. Resemblance is in the eye of the beholder.)
So what should adobe pay you for using the data in training? Some kind of fraction of the overall revenue they generate from the new product? The license currently used for their stock program make it seem like they don't have to pay anything at all, because this use cases wasn't understood previously. Adobe reserved to rights to do it, so legally they can - but if they want to continue getting contributions they will need to figure out some kind of updated royalty sharing agreement.
That is the question we are asking, yes. Based on the reading of the contributor agreement it sounds like Adobe doesn’t have to pay a cent to the creators to train models on their work.
Does that sound fair to you?
When the agreement was signed no one was even able to imagine their work being used for AI. As far as they knew they were signing a standard distribution agreement with one particular rights outlet, while reserving all other rights for more general use. If anyone had asked about automated use in AI it's very likely the answer would have been a clear "No."
It's predatory and very possibly unlawful to assume the original agreement wording grants that right automatically.
The existence of contract wording does not automatically imply the validity of that wording. Contracts can always be ruled excessive, predatory, and unlawful no matter what they say or who signed them.
Maybe. Maybe not. Very clearly there is a price point where it could be worth it for the artist. Like if adobe paid more for the rights than they recon they will ever earn in a lifetime or something. But clearly everybody would have said “no” at the great price point of 0 dollars.
Does that sound fair to you?
See how stupid that sounds?
A tool maker does not have a claim on the work made with a tool, except by (exceptionally rare) prior agreement.
Creative copyright explicitly does give creators a claim on derivative work made using their creative output.
That includes patents. If you use a computer protected by patents to create new items which specifically ignore those patents, see how far that gets you.
I expect you find this inconvenient, but it's how it works.
Like, if you created a lovely piece of art, hung it on the outside of your house, and I was walking on the sidewalk and viewed it. I would not owe you money and you would have no claim of copyright against me.
Copyright covers copying. Not viewing.
So an AI views your art, classifies it, does whatever magic it does to turn your art into a matrix of numbers. The model doesn't contain a copy of your art.
Of course, a court needs to decide this. But I can't see how allowing an AI model to view a picture constitutes making an illegal copy.
Memory involves making a copy, and copies anywhere except in the human brain are within the scope of copyright (but may fall into exceptions like Fair Use.)
Supposing the AI said "here's the picture you asked for, btw it's influenced most by these four artists and here's links to their works on Adobe Stock"... is that better (or worse)?
No actually, not for this situation. They don't if they sold the right to do that, which they did.
> except by (exceptionally rare) prior agreement
Oh ok. So then, if in situation 1, and situation 2, there is the same exact prior agreement on the specific topic of if you are allowed to make derivative works, then the situations are exactly same.
Which is the situation.
So yes, the situations are the same, because of the same prior agreement.
Thats why the situation is stupid. The creator sold the rights to make derivative works away. Just like if someone sold you a computer.
And then people used the computer, and also used the sold rights to make derivatives works for the art, because both the computer and the right to make derivative works were equally sold.
> which specifically ignore those patents
Ok now imagine someone sells the rights to use the patent in any way that they want, and then you come along and say "Well, can you considered that if the person didn't sell the patent, that this would be illegal?"
That wouldn't make any sense to say that.
It's completely different than many (most?) other companies, which are training on data they don't have the right to re-distribute.
I think you are making a jump here. I’m not a lawyer, but your first sentence seems to be about why it is legal. And then you conclude that that is why it is also fair. I’m with you on the first one, but not sure on the second.
The creators uploaded their images so adobe can sell licences for them and they get a share of the licence fees. Just a year ago if you asked almost any people what “using the images to develop new products and services” mean they would have told you something like these examples: Adobe can use the images in internal mockups if they are developing a new ipad app to sell the licences, or perhaps a new website where you can order a t-shirt print of them.
The real test of fairness I think is to imagine what would have happened if Adobe ring the doorbell of any of the creators and asked them if they can use their images to copy their unique style to generate new images. Probably most creators would have agreed on a price. Maybe a few thousand dollars? Maybe a few million? Do you think many would have agreed to do it for zero dollars? If not, then how could that be fair?
Have you read the contributor agreement? That seems to contradict what you are saying.
Any country that tries to forbid this will be torpedoing their economic competitiveness against countries that allow it.
I don’t see why you are asking this. Which part of my comment made you think it is preferable?
And that's a different question from whether or not they deserve extra compensation. Is it moral or ethical to use their work to directly undercut them via ai 'copying' their work?
Half the problems with music is because of record companies magically inventing new ways to try and extract more money from each other and their supply chain.
"Oh, your band looked at some hookers they passed on the way to the recording studio? Well, they obviously owe those hookers a cut of the royalties now for inspiration..."
Trying to use AI as an excuse to be paid a 2nd time (for previously fully paid works) seems like another attempt at rent seeking in a similar manner.
Adapting what Steve Albini said at the time: If you're an artist, negotiate your future "AI training rights" separately. Make sure they're in writing. Make sure attribution rights stay yours.
It merits investigation as to have these creators been “fully” paid to the extent that they have no claim to any future royalties and can have no objection to their work being used as training data.
There are also a few other more Adobe-specific restrictions.
I should have been clearer and specified sale of rights rather than licensing of a stock image in my original reply, which I now realize is confusing.
Companies certainly can offer more narrow terms/commitments about what the stock images can or cannot be used for if they want to, but economically they're incentivized to maximize their own freedom, and not to negotiate a complex bundle of rights for individual images. Of course for journalistic and artistic images that are unique in some way, the original creator can negotiate.
What a time to be alive.
I instantly thought of how bitchin' their library of images must be.
Can you tell us how many images/size of set?
Fine tuning Stable Diffusion with your own images is way easier than creating Stable Diffusion in the first place.
If you're creating your own I stand corrected and that's some serious investment.
But training your own is pretty doable if you have the budget and enough image/text pairs. Most people don't have the budget, but at least Midjourney and Google have their own models.
(i work for Adobe)