High Court rules that Getty vs. Stability AI case can proceed
technollama.co.uk
technollama.co.uk
Bizarre.
Can you imagine if you had no piracy occurring in the US when users download pirated content from a server in the Philippines and you could only handle the case under the Filipino legal system?
There will be a separate trade mark question on the outputs, which wasn't part of the request to dismiss the case.
>Can you imagine if you had no piracy occurring in the US when users download pirated content from a server in the Philippines and you could only handle the case under the Filipino legal system?
Your hypothetical is about actual distribution and communication to the public of copyright works, this is very different to what is at stake here.
You could generate a billion images with SD and train the next model on them. Make sure they don't look close to copyrighted works. Being AI generated, they have no copyright. You can still use real data as well if it is in the public domain.
If you do this enough the initial copyrighted dataset is going to be further removed from the model. The model can't reproduce a copyrighted work because it hasn't seen any of them during training.
But more importantly, this process strictly separates ideas from expression and trains only on ideas without copyrighted expression. If authors complain it means they want to own ideas and styles.
You can also use copyrighted works to train a classifier to rank the quality of training examples, and apply it to filter your synthetic data to be higher quality.
You can even train a RLHF model to say when two works are "close enough" to constitute an infringement, and double down on safety by ensuring you don't generate or use risky works.
That's why I was saying that I don't think copyright has much meaning left in it anymore. Knowledge wants to be free, it travels, shape-shifts and evolves. It does not belong to any one of us except if we keep it to ourselves.
But you can also empower the LLM, it can use more tokens, chain of thought, multiple rounds of LLM inference, use other specialized models, use tools, execute code, use search, and have a human in the loop. Does that remind you of anything? Yes, "OpenAI GPTs", they are the empowered LLM environments that can create data superior to LLM alone.
The general recipe is LLM+something extra = smarter than LLM. That something extra is usually a simulation or code execution or some real world interaction. This generates LLM error examples and feedback to learn to fix them in the next iteration. It is targeted training data.
Knowledge is easy to scrape, transform and train on. Once it gets into the open datasets, all models trained on that data will inherit the skills.
In language, more recently the most useful training data has been generated with GPT4. Its abilities get transferred very efficiently to lower models.
I think the same will happen with images. We're going to generate synthetic datasets that would replace the original copyrighted works, but be more efficient and diverse.
The concept of having copyright over a style is ridiculous anyway. Our courts should have their time taken up by making subjective determinations of whether some art is in the style of another.
If new tech circumvents the law, the law can easily be changed, so:
> If authors complain it means they want to own ideas and styles.
is "yes, and?"
Also, copyright as it currently exists (a legal construct) is supposedly to promote the arts. GenAI may obviate the economic need to promote the arts… but if art is to us as a peacock's tail, then it won't ever obviate the desire to promote (and protect) the arts. The expense of human labour may be the point.
It would be self defeating for artists. The same tools developed to scan AI art for copyright infringement will trigger on their own works as well. All "side inspiration" will be revealed, it will have a chilling effect on freedom.
Do you want art to be made of little islands of copyright, where someone staked their claim - nobody else is allowed to create? If they can own styles they can ban others from using those styles.
Let's explain it like this: in a system, a user has ownership over a few files. But from now on users will own whole extensions, like ".json" and own everything "*.json" instead of "myreport.json". They want copyright wildcard "*.©", it's a power grab. They want copyright to be more like patents.
Perhaps; frightened and angry people often can't see the consequences of the direction they're running in.
But: if the effort is the point of art, then "proof of work" is how that goes down, not your scenario. Filming the artist as they put oil to canvas etc.
> They want copyright to be more like patents.
Not patents, trademarks. Patents cover novel specific inventions for a few years, trademarks cover anything that's similar enough it might confuse a customer for as long as you maintain it.
And what I want is almost irrelevant, the question is the aggregate will of society. My preferences feel like they're negatively correlated with public opinion on the topics loud people discuss most.
Careful with this assumption. The US Copyright Office have said this. Some other jurisdictions may have similar statements. But few (any?) claims around this have been tested in court, and specifically there is still the risk that there may be jurisdictions that find the output to be infringing if the training data is.
It will take time for this to shake out.
I do mostly agree with you that I think enough doors will be left open that it will be possible to find paths that will effectively "work around" copyright in fairly significant ways, but it's not so open and shut what will be ok and what won't.
I think if you replaced human with computer and went through all the training, you come to the conclusion that the human is recreating copyrighted works . so if you the send that human out to all these people knowing what they know, they're creating infringement.
the hinge will be that the training data was never licensed by the model maker, the trainer for the human computer analog.
the only thing limiting damage is that the scale of impact is unmanageable.
What an outcome like that would do, however, would be to massively centralise the ability to train legal models.
At the same time you'd end up with a lot of "washing" of training data (a lot of "works for hire" that'd really be outputs of models trained on copyrighted works), and a lot of development moving to other countries.
I'm pretty sure we are at the low hanging fruit and waiting for a mechanical turk to properly put a training set together.
Seems like jurisdiction would be based on the copyright of the allegedly infringed images, and UK-based users creating copyright infringing copies in the UK.
But that's apparently not the law or case law in the UK yet.
There’s no question these neural networks and their output are derivative works. However being a derivative work isn’t enough to guarantee copyright infringement.
So, the only question is if we are going to carve out an exception here or not. The idea someone can use a VCR to copy live TV and let people watch it later came out of a court case not copyright law. There’s a lot of such exceptions, but getting one isn’t guaranteed.
The big one is recipes. Recipes under the current copyright regime in the US are considered non-copyrightable facts, which is why every cookbook and recipe blog has lots of copyrightable splash photos and personal anecdotes. Congress specifically doesn’t want grandmas getting sued for copying the recipe on the box.
fair use is mostly a US concept, there is no such thing in the UK or most other countries
https://www.twobirds.com/en/insights/2020/uk/intellectual-pr...
https://www.copyright.eu/docs/protection-of-a-recipe/
Though you can patent novel methods of food production, which is also true in the US.
The root statement is still the same, legislatures can amend copyright laws as they wish if they really care. I don’t know that the UK parliament is exactly functioning well right now, but that’s my impression from across the pond.
in terms of ability to legislate it works considerably better than the US congress
up to you if you call that well functioning
Recipes don't have a specific exception within the the copyright law that Congress has carved out.
It is also not cut and dry. It basically boils down to facts not being copyrightable. So a list of ingredients and basic instructions (e.g. cooking time and temperature) won't be granted copyright protection.
But, the prose in the instructions can be copyrighted. So copying a whole recipe verbatim can be copyright infringement, but copying the list of ingredients and writing out the basic instructions is not.
Most generated content almost certainly isn’t derivative work by the standards of copyright law. It’s plainly obvious to anybody who’s read Frank Herbert’s books that he derived a lot of ideas from Isaac Asimov, but it’s equally obvious that Dune isn’t a derivative work of Foundation.
If I had some commercial interest in generative AI models, I’d be very happy that everybody is debating the copyright implications. Because copyright law is certainly going to favour the models. The biggest regulatory risk to them as far as I can tell is that they clearly don’t have section 230 protections, and I can’t imagine how that isn’t going to come crashing down around them rather soon.
Specific examples of clear copyright infringement mean that output is a derivative work AND by encoding enough information to recreate it the underlying neural network must itself be a derivative work.
In the two US cases we have any progress on so far, the established requirement for substantial similarity (opposed to "dependant on" or such) has been upheld, with Judge Vince Chhabria specifically setting out that it'd "have to mean that if you put the Llama language model next to Sarah Silverman's book, you would say they're similar". and Judge William H. Orrick agreeing with the defendants that "plaintiffs cannot plausibly allege the Output Images are substantially similar or re-present protected aspects of copyrighted Training Images, especially in light of plaintiffs’ admission that Output Images are unlikely to look like the Training Images".
The UK definition of derivative works is, to my understanding, narrower and specifically enumerated as opposed to the US's more open-ended definition.
The remaining area of doubt, assuming the above remains consistent, is over the transient copying that occurs during training.
i think this should be dismissed as it is the same level of transience as the workings of the internet; you and your ISP, caching proxies etc, all made a transient copy as part of the existing (legal) consumption of the works that the author has put online.
Unless the works was illegally copied for training - which cannot be true if the works was publicly available for viewing on the internet, this transient copying cannot be a valid infringement.
Downloading a singe transient copy of some image once in the lifetime of a company is different than doing that same action a hundred times once for each version of the network.
Hard to argue keeping a copy of some copyrighted work indefinitely counts as transient.
Precisely to argue for transient copies, they don't need to keep terabytes of data stored.
>Hard to argue keeping a copy of some copyrighted work indefinitely counts as transient.
You're assuming that they're keeping the works indefinitely, which again is not the case.
Those kinds of legal workarounds rarely work.
They are dependent persistent access allowing them the equivalent benefit of keeping a persistent copy.
Defendants can easily argue that being 1/10 millionth or whatever of the training set means their specific work is unlikely to show up in any specific example but the underlying mechanism means it can be recreated.
A derivative work is an expressive creation that includes major copyrightable elements of a first, previously created original work (the underlying work).
There is absolutely no agreement that what neural networks do (as a rule) counts as such, so it is not at all correct to say "there is no question..."
If learning how to draw by watching other people draw makes everything you draw a derivative work, then perhaps you have a point.
For a neural network to be able to recreate a complex work with minimal prompting it must be encode that information and therefore be a derivative work.
Judge Orrick in one of the US cases already called this idea 'nonsense", his words.
Specific and clear examples of derivative works are shown therefore both those exact examples and the underlying neural network must be a derivative work.
So if they decides to play this game then so should we.
Using publicly available information doesn’t require anybody’s consent.
Are text snippets, thumbnnails and site caches shown by search engines (on an opt-out basis) "theft"? If you draw a car, which you can do due to having seen many individually-copyrighted car designs, are you stealing from auto manufacturers? Have I just committed theft by using a portion of your comment above as a quote?
I don't claim here that statistical model fitting inherently needs to be treated the same as the above examples, but rather use examples to show that the bar of "using" is far too broad.
Legally, copyright infringement in the US requires that the works are substantially similar and not covered by Fair Use. Morally, I believe that artificial scarcity, such as evergreening of medical patents, is detrimental and needs to be prevented wherever feasible - and wouldn't call any kind of copying/sharing/piracy "theft". The digital equivalent of theft is, for example, account theft where you're actually removing the object from the owner's possession.
That's a perfect oxymoron.
It does not assume that.
Otherwise there's obviously a legally relevant distinction between a human mind which is ascribed agency to decide if and how to use its memories of copyrighted material, and importing into an information retrieval system which can't help but spit out transformations of parts of its inputs on demand, (including lossy representations of the Getty watermark if it's fed enough Getty material, or an exact facsimile of an image if that's all it's trained on...)
Copyright uses infringement, which is not theft: it's non-rivalrous, and it contains a number of exceptions.
Someone needs to come up with a royalty structure for this stuff.
Even the argument that its logical to call it that isn't certain.
If i take your picture I own it?
If you take a picture and i upload it to fb meta gets to use it?
If I publish your book under my name and no one finds out, did I write it?
If no author can be found, may I read it?
no you wouldn't, but these diffusion models do way more than ffmpeg, and do qualitatively different things.
I am on the fence, but i lean towards the side where training an AI using existing works is not infringement, as long as the AI's output is (or can be) majority new works. For example, a poor training algorithm that merely repeats the training dataset (and cannot output new works) is infringing, while a different algorithm (such as the current stable diffusion one) that can output works that has never been made and is totally new, does not infringe - after all, style and ideas are not infringing and if the algorithm managed to extract those ideas from the training set, all the better.
If I look at a photo of Prince and then using that image as reference create a new silkscreen painting is that fair use or infringement?
Because the US Supreme Court has ruled that instance I referenced was infringement as both images were used for magazine covers [0].
the existing copyright rulings are sufficient to determine this, and has nothing to do with ai models.
You've already pointed out a case - if you use an AI to generate an image which has sufficient likeness to an existing one, then the AI portion is irrelevant to the ruling. You could've made that same image in photoshop without AI, and should obtain the same ruling.
But in the above circumstance, the silkscreen used in the creation of the image does not itself infringe. And replace that silkscreen with AI model, nothing has changed.
so by that standard, why isnt photoshop a copyright infringement? You can use it to create a copy just the same.
But it isn’t. It’s just a series of vectors that point to a likely occurrence of the next word or pixel or bit in a sequence.
https://twitter.com/StefanKarpinski/status/14109710611816816...
It’s hard for people to understand this concept, but the fact that a model repeated some data verbatim is a happy coincidence (!) solely based on patterns of data that it seen before.
I think people have also have a hard time with how these models are trained. They are vacuuming up all sorts of data and learning from them by creating vectors that determine how follow-up data should be generated.
Sure, the original creators of this content aren’t being compensated or even recognized for it. I don’t have a good idea on how that should be handled.
For normal humans though, looking at art or reading a book, and later repeating some passage or drawing something from your own memory is not a crime. (Unless you’re sharing the DeCSS source code I guess…)
Slightly changing the topic here, but I do wonder what were to happen if someone wrote a program called “Monkeys on Typewriters” that just iterated through various combinations of characters (or bits or pixels) and was able to recreate things verbatim.
Is that random happenstance copyright infringement?
False, actually; memorizing a copyrighted work and reproducing it other than in conditions specifically excepted from copyright protection is a violation of the exclusive rights of the copyright holder to make copies.
Copyright doesn't just apply to mechanical copies which don't have a human brain in the middle of the process.
That being said, Getty is hardly the paragon of goodwill considering they regularly steal from public domain databases, issue DMCA takedown requests of the stolen content from said databases, and then turn around to sell it to unwitting people for a subscription. They own none of the copyrights for what they are doing but have been allowed to get away with it.
What’s the legal distinction between you learning and AI learning?
I'm not distributing my brain, at least same (but probably more restrictive) should apply to models - training is okay, but using and distributing should be limited by copyright
Besides which, "learning" isn't a fair use exemption anyway.
There are public domain works you can use and copyright doesn't protect ideas. It protects expression of ideas, so getting "just the ideas" without the expression is ok.
The problem is that "expression of ideas" in the realm of AI is akin to plagiarism by human standards, because its a literal copying of the source material blended together. I couldn't recite you the entire plot of the Odyssey off the top of my head literally, but AI can, because it has the source material. We just tell it to do funny ha-ha things so its okay.
If the license makes the data public viewing, like with websites, then slurp all you want. If the license forbids automated bulk processing, then stop whining about fair use and pay for a license that allows bulk processing.
"Out system actively uses every single byte of data to produce any output" is so obviously not the intention fair use clauses.
Either you insist that copyright must be respected at every level, and the creators of material used for training deserve appropriate compensation, or
You throw out copyright completely in this context, but that means the resulting models cannot be treated as proprietary either unless they were produced using absolutely no unlicensed training data.
I think there is an argument for both. Want to create a proprietary model for commercial use? Pay up. Creating an open source, copyleft project exclusively for personal use and artistic expression? Exemption.
The current status quo is perfectly described by powerful corpos extracting rent. Billions for themselves and pennies for the average artist.
I don't think that current copyright laws automatically entitles people to royalties from something like AI-generated imagery. The dichotomy you've presented here isn't pro-copyright vs anti-copyright, but "so pro-copyright that they argue for expanding the current laws" vs not.
> Want to create a proprietary model for commercial use? Pay up. Creating an open source, copyleft project exclusively for personal use and artistic expression? Exemption.
That definitely benefits all the "powerful corpos" you've mentioned here. Now, Disney, Adobe, Meta etc. can use a fraction of their money to get all the data they would ever need and be the sole profiteers, while all newcomers will face an impassable barrier to entry that prevents them from ever threatening the existing players.
This would mean we have to do a few difficult and worthwhile things: explicitly dismantle the copyright system, encourage artists to donate their existing works to the commons, and then only make datasets based on legally collected information. This would also have the side effect of encouraging the development of new training techniques and model designs which are more sample efficient.
I am afraid that what we will do instead is allow some erosion of copyright for small creators without dismantling the power large intellectual property holders have over the rest of us.
I think "take" is the wrong word here, nobody is republishing the copyrighted works, instead the model gets a gradient update. The update is shaped exactly like the model itself, and it gets stacked up with other updates from other examples. It doesn't look like the original work at all, the original work was a picture or book, the gradients look like a set of floating point tensors. AI models decompose inputs into basic concepts, they don't copy like bittorrent.
Why should an AI not be allowed to form a full world model that includes all published works? It's not like the authors can use copyright to stop anyone from seeing their works, they never had a right to stop others from seeing.
Whether or not it is taking is more nuanced, but I will say I’m not sympathetic to the idea that it’s broadly similar to a human looking at the work. It’s just very, very different. You can’t spin up a copy of a human on a cloud server and make them work 24/7.
I would expect that as laypeople we aren’t equipped to reason about this effectively. I suspect that decades or more of case law would be relevant to how this would be viewed, and I’m personally not equipped to argue it.
What I do know is that artists don’t feel good about it. They feel like they’re being taken from. And I’m not inclined to quickly dismiss their concerns. I think this needs careful, deliberate consideration. And if a system could be built that is consent based, I’d feel much better about it. A human child could be raised and mature without ever being exposed to copyrighted material beyond a handful of books (harder in the modern world but common 200 years ago). Maybe we just need to build better models. It certainly seems possible.
The only limitation that needs to be there for training on copyrighted works not to be infringing is to accept that extracting information about the work is not infringing if copyrighted elements of the work itself is not significantly reproduced.
It is interesting to me that we are finally seeing a case where many smaller, independent artists and creators are using these laws to assert their rights against the encroachment of the moneyed tech interests, and of course now all of the powerful corpos are singing another tune. Rules for thee, not for me.
> This is no longer possible as far as I can tell.
? Why would it no longer be possible?
Stability has exactly zero way of updating the weights once they have made them public. Is the suggestion that this was only possible on 1.4 and 1.5 and XL don't have the issue?
...or somehow what the model was previously capable of doing is now no longer possible?
That seems like enormously unsubstantiated speculation.
We have proof from multiple independent studies that these models in general memorize a small percentage of their training data to the degree that reasonable reconstructions of the original training images can be recovered from the model.
There is, to my knowledge, no mitigation of this that has been implemented, or even conceptualized by either stability or anyone else.
Interesting read regarding the actual ruling the judge made, but this is some "opinion here" commentary, which seem out of place; what the court actually rules is the interesting part of this.
I think we need an entirely new approach to copyright. You can't copyright a brain, and weights are closer to that than they are pirated digital archives.
I don't understand why I keep seeing people roll stuff out like 'should I be lobotomized' as some sort of attempt to distract from this just being a dispute over copyright minimalism/maximalism.
If it produced images identical or highly similar to the copyrighted work, it would be similar to you making a copy and giving it to someone without paying royalties.
The article also shows an example of it reproducing their trademark.
Edit: Rephrased to be more clear.
I’m saying; if you’re covering a court case, don’t go off on a tangent and start saying things which are misleading or false with regard to the case.
> I think…
This is precisely my point. Most people don’t care what I think; or, you, or the OP on this topic.
Opinions are thick and fast.
What people care about is what the court rules.
This adds the caveat "as far as I can tell". That's a caveat that the capability to produce Getty outputs has been removed, it can be tested and amended if new information comes to light.
>What people care about is what the court rules.
Courts also rely on experts and expert testimony, the judge here is even citing a law professor's opinion. The idea that judges never rely on legal experts and commentators is strange.
Then it is creative. It is doing meaningful work on top of the basic facts it learned. The facts and the style can't be copyrighted.
Can I download you via a URL?
If not this analogy is completely meaningless. You and stable diffusion are just too different.
You can download an LLM trained on my writings. And I have a speech model, too.
It won't be much longer until we have brain scans of people that are turned into some form of generative system.
I've tried with newer models, and you can't produce outputs with Getty logos, if you use "getty images" in prompts it won't work. Even if you type "Getty" you don't get the logo, I remember reading a tweet by Emad that any reference to Getty had been removed as of 1.5, but I can't find it anymore, maybe deleted now?
So yes, this is a thing, feel free to try it for yourself.
Emad
They went on to sue Google Images that allowed the spirit of open web - where you could simply right click and save any image - from doing so.
And now that AI can pretty much produce stellar quality images with just a single line of text input, they don't have anything else to do other than drag everyone into litigation.
I'm a photographer myself (Sony A7C2/24-70mm G-Master) and I have learned to accept that this is going to be the future. And though there will always be a market for real photographers, AI will take over bulk of our jobs. And the programmer in me who spent ages trying to find the perfect background image for my website says that's not really a bad thing.
If the bit sequence is likely to occur because it's someone else's creative content (or part of it is)... that doesn't seem like it can be a 'fact' in the relevant manner.
I agree that a particular sequence of words is copyrightable.
What I'm struggling with is that facts _about_ that corpus of text are not copyrightable. A simple fact could be that the word "bar" is the 5th word. The 6th word is "jazz". Etc.
A model is trained from these "facts" across many source documents. It is thus itself a derived 'fact' given a set of training inputs and parameters, so then how could _that_ then be copyrighted?
Put another way - there's the origin text and then.. is it turtles all the way down and none of it can be copyrighted because its all math and calculations derived from that?
The fair use excuses here are absolutely weak in Stability's case and this will only end with a licensing deal being made.
There's no fair use in the UK, and this decision is preliminary on whether the case should proceed.