Nvidia is sued by authors over AI use of copyrighted works
reuters.com
reuters.com
>AI's rise has made Nvidia a favorite of investors.
>The Santa Clara, California-based chipmaker's stock price has risen almost 600% since the end of 2022, giving Nvidia a market value of nearly $2.2 trillion.
Me and Joe Nobody wouldn't be able to afford it, but I find it hard to believe a trillion-dollar company couldn't get proper rights to the data it uses to train its models.
Does it?
And if it did, why wouldn't fair use apply here as much as it would apply to the actual training/SaaS offering?
Fair use generally only applies for parts of a work, and not using the entire thing. AI companies are claiming fair use, claiming that the processing is transformative. This needs to be determined in court. I think the court cases currently underway will show it is not transformative, because the original work can be recovered. Fair use is also determined on a case by case basis, so I think courts won't have the power to grant blanket permission. New laws would need to be passed for that.
Because of the substantial non-infringing uses. A court wouldn't throw away a model that can do non-infringing stuff 90% of the time if it can infringe copyright the other 10%.
> Fair use generally only applies for parts of a work, and not using the entire thing. AI companies are claiming fair use, claiming that the processing is transformative. This needs to be determined in court. I think the court cases currently underway will show it is not transformative, because the original work can be recovered. Fair use is also determined on a case by case basis, so I think courts won't have the power to grant blanket permission. New laws would need to be passed for that.
Fair use can apply to any possible allegation of infringement, not just parts of a work, as long as courts find the facts are in favor of it.
Whiile fair use is a case-by-case thing, courts can set precedent that will apply to all situations with similar fact patterns.
https://spectrum.ieee.org/amp/midjourney-copyright-266687210...
No, it absolutely does not.
There's not enough space in even the larger models to contain verbatim copies of the training data, or anything remotely close to it.
Reminder: anyone nerd enough to be on this site is on the top 5% of computer skill, and 20% of people in the planet can't even use a computer, and 30% that can wouldn't be able to perform a simple multi-filter task such as "find a sustainability-related document that was sent to you by John Smith in October last year."[1]
I believe if the web is meant to be open, it can not antagonize 50% of the people in the world who can't understand technology. It's just not in the spirit of the web to say "don't put it on the internet." The web was made to put things on the internet. When we can't do that anymore, the web is dead.
We're going to be in an awkward situation where pirating can net you a $150000 per mp3 fine but big companies can train on the said mp3s and produce a nearly identical product for you.
But that's why we have a giant police and propaganda machine.
Although it would be hilarious to think what would happen if the news industry went to war for instance, feeding false information to the AI scrapers making them look ridiculous.
I think it also ties into not wanting AI to be trained on AI generated content. The winners will want to license original sources, and licenses work both ways so you can sue if paid for clean data and got dirty.
What? Google Books won.
I kind of feel that in general with AI regulation, the USA will have to go all-in somehow. EU will delve into future obscurity, but this technology is too valuable and China is breathing behind's USA neck, so I expect law to continue to be lax around IP of data used to train AI models.
It would be best if a non-profit, the equivalent of EFF or FSF could take on the trolls in the name of software freedom and ideally get some legal clarity, as opposed to just getting some deals struck with big companies.
Alas...
And if copyright applies, wouldnt it imply that someone learning something from a book could also then be controlled by the licensing of said book on how their gained knowledge could be utilized in the future?
Obviously the answer here is for companies producing AI to curate, obtain, and/or pay for fully legal training data. The problems have been that gathering and using copyrighted data is very easy, and AI is extremely data hungry (some experts theorize massive data alone is responsible for AI success, and algorithms are secondary at best), and there’s at least the perception if not the reality of a high stakes winner-takes-all race to produce the best AI.
To me this feels a little like the situation tech companies have put themselves in with automated support and no way to reach humans, in that they couldn’t have scaled like they did and gotten there without dropping hands-on support on the floor, but they’ve created a time bomb that is beginning to backfire in more and more serious ways.
The minute copyright lawsuits begin to land, suddenly the copyright becomes the moat. Big companies will license, and everyone else will wither on the vine. Google and Facebook will reign supreme.
I'm terrified by that. I hope we have a few more years to compile open data sets and tools.
[1] That is to say, thus far, product is the only moat, not models and not data. Most companies have razor thin product so far.
software - no idea. It won't be banned for sure, but something like 5% royalty, could happen.
In any case, exact reproduction of the full novels as part of a dataset is of course not covered by fair use and proving that through discovery will be trivial.
It's an admission to copying when there's a prima facie case of copyright infringement, but wouldn't generally be considered actual infringement if the fair use defense is successful. It's perfectly fine to rely on fair use - we'll have to wait and see how it applies to machine learning, but it's not some obscure implausible theory.
> In any case, exact reproduction of the full novels as part of a dataset is of course not covered by fair use
It definitely can be. Consider in Authors Guild v. Google, where Google copied millions of in-copyright books, stored them internally in full, and made them accessible as snippets: "verbatim intermediate copying has consistently been upheld as fair use if the copy is ‘not reveal[ed] . . . to the public.’"
No, it is not. It's an admission of copying, but not an admission of copyright violation. Legal copying (which includes fair use) is not a "violation" of any kind.
For example, in general it is unlawful to use controlled substances (e.g., morphine).
However, if one has a prescription for morphine and uses it in accordance with that prescription (i.e., legally), that isn't an "admission" of "violating" the laws against narcotics use.
> If this is your starting point, you should probably look for new lawyers.
Someone needs new lawyers, but it's not the OP.
You know what's the most popular <style> for Stable Diffusion? Photorealism. And the second most popular <style>? Anime. At the risk of stating the obvious, I will note that no artist can hold the honor of having invented generic "photography" and generic "anime".
But sure, I will give you 2 internet points for accurately spotting one of the extremely rare use cases of Stable Diffusion where an artists' style is copied. I agree with you that is not all kosher.
> I mean the whole point of using AI at all is to get something to give you high quality stuff for free, where the high quality has been lifted from a high quality training set.
No, it's not. It's a failure mode. It's called "overfitting". We went over this.
The vast majority of Stable Diffusion's usage is about producing something novel, and reproducing something that looks like a training image is considered a failure.
That’s entirely beside the point. Copyright still applies to photographs, and training has been occurring on copyright-protected photographs.
> It’s a failure mode
It seems like maybe you didn’t fully parse what I said there, which wasn’t about reproducing works closely enough to violate copyright, it was about the broad goal of AI, which absolutely is to train on fixations (in Copyright parlance) and be able to produce mashup results that mimic the inputs as closely as possible without breaking the law. Because the goals here and the design intent here are to skirt the legal lines, it’s not at all surprising that sometimes we will cross it, whether it counts specifically as overfitting or not. When the broad design goal combined with users and prompts seeking specific outcomes is riding the line between mimicking style and clear copying, we’re basically priming the overall system to live in a legal gray area.
This is outside the question of whether consuming copyrighted training material is legal, regardless of whether any reproductions occur. Watching copyrighted movies is not legal, I don’t have to share them with anyone else. Why should AI training that consumes copyrighted material even be legal in the first place? (Maybe it’s not.)
No, you missed the point. I was making a distinction between artist styles, like "in the style of Picasso", versus generic styles, like "photography". Furthermore, I was making the point that the generic styles, such as "photography", are way more popular in the Stable Diffusion community, compared to artist specific styles, like "in the style of Picasso".
Just to spell it out as clearly as possible: the majority of Stable Diffusion usage is not trying to imitate the style of a specific artist.
Whether or not training on copyright-protected photographs is legal is tangential to the point.
> It seems like maybe you didn’t fully parse what I said there, which wasn’t about reproducing works closely enough to violate copyright, it was about the broad goal of AI, which absolutely is to train on fixations (in Copyright parlance) and be able to produce mashup results that mimic the inputs as closely as possible without breaking the law.
I fully understood your point and expressed my disagreement of it. Specifically, what you described is typically not the intended outcome, it is a failure mode. I have repeated this many times now and you keep ignoring me, as if the problem goes away by ignoring it.
As far as copyright is concerned, there is no such “generic” category of photography, you’re trying to invent a term that has no legal status. All photos fall under copyright. Training on them may break copyright, reproducing them accidentally may break copyright.
It is irrelevant what the majority use case is. (But how can you possibly claim to know what the “majority” of Stable Diffusion’s usage is?) It is a fact that examples of people asking to mimic specific artists are all over the internet, it’s quite common and easy to find. It doesn’t matter if this usage is not the majority, all that matters in terms of the legal question is whether the usage of AI meets the criteria established in the law.
You’re splitting a hair trying to contradict me when I’m trying to point out that regardless of whether people are trying to imitate a single artist, the design and implementation of the system is expressly for the purpose of imitating multiple artists. My point doesn’t hinge on single vs 2, 3, or many. This isn’t a failure mode, this is the stated purpose of many of today’s neural networks. They are built to copy snippets of other people’s work.
Not related to what I was saying at all. My point is that producing an AI image "in the style of a photograph" is less of an infringement on artists' rights than producing an AI image "in the style of Pablo Picasso". Surely you can agree with me on that point?
Whether or not the AI models' training on copyrighted photographs was legal or not is tangential to the point above.
> But how can you possibly claim to know what the “majority” of Stable Diffusion’s usage is?
I'm actively following what people create, share, and like on Discords, subreddits and Civitai. I would say less than 1% of popular works are recreating the style of a specific artist.
> It is irrelevant what the majority use case is.
Well, you founded your whole argument on the very specific niche case of "create an image in style of <artist>". If that's now irrelevant, then I don't know what your argument is anymore. Maybe you shouldn't have made that argument in the first place if you find it so irrelevant.
> It is a fact that examples of people asking to mimic specific artists are all over the internet, it’s quite common and easy to find.
Sure. It's also a fact that examples of people killing each other with knives is easy to find. That doesn't mean the majority use case for knives is killing people, and that therefore knives should be banned.
<thing> can be used for <good> and <bad> and we have to look at both before we say "ban <thing>", it's not enough to just look at <bad>
> You’re splitting a hair trying to contradict me when I’m trying to point out that regardless of whether people are trying to imitate a single artist, the design and implementation of the system is expressly for the purpose of imitating multiple artists. My point doesn’t hinge on single vs 2, 3, or many.
I'm not splitting hairs on whether an image is imitating "style of <artist1>" or imitating "style of <artist1> and <artist2>". A distinction like that is not meaningful. If you imitate artist1 and artist2 in the same image, that's still imitation of specific artist styles.
Let's take an example. Here is the top image on Civitai today: https://civitai.com/images/7511111
That image seems to be using a LoRA that's intended to reproduce "China-Splashed Ink" style. I have personally seen this style in various artworks, movies, and games. For example, the Total War: Three Kingdoms game uses this style. This is a style that a large number of humans have contributed to over a long period of time. There isn't any individual artist who can claim that it's "their" style.
Do you have an issue with AI image generator generating an image in that style?
It depends entirely on what image was produced. The prompt isn’t the infringement, the output is. And photographs are no less copyrightable than any other art. So I can’t agree with your statement as written because the statement doesn’t make sense and is not true. If an AI’s output happens to look like one or two specific photos, then it violates copyright, regardless of what the prompt was. (The photo doesn’t even have to be in the training set!)
You might have gotten stuck or confused on the idea of whether or not users actually prompt for a specific artist or art style. I stopped talking about a prompt to copy a named artist several comments ago and started pointing out that AI inference can break copyright without a prompt to break copyright.
So, If you’re trying to give a hypothetical example of AI producing an image that looks like a photograph, but not one that looks like any existing photograph, then I can agree that this inference doesn’t violate copyright. But if the output does end up looking like an existing photograph for any reason, then it does violate copyright. This doesn’t depend on whether we’re talking photography or drawings, the standard is the same regardless of the medium. The balance of photos in the training set might change the likelihood of accidentally reproducing one of them, and that’s an interesting question, but not what we’ve been discussing.
> I'm not splitting hairs on whether an image is imitating "style of <artist1>" or imitating "style of <artist1> and <artist2>". A distinction like that is not meaningful. If you imitate artist1 and artist2 in the same image, that's still imitation of specific artist styles.
Good. You’re getting closer to what I’ve been trying to say. AI always imitates the style of some number of specific artists, because it is always trained on some number of specific artists, and it usually inferences from fewer than all the artists in the training set. The number might be high for any given run, and a particular style might not always be identifiable, but the design of the system is, in fact, to copy from multiple artists at all times. I agree with you that the distinction between one and many is not that meaningful, however in the courts there is precedent for declaring a copy of one artist to be a violation of the law while letting a remix of several artists off the hook. From the copyright perspective, AI inference might not violate copyrights as long as there is no identifiable infringement of a singular work.
Whether or not training on copyrighted materials in the first place infringes has not yet been litigated, but since humans get threatened and sued for it all the time, there’s reason to believe AI could have legal trouble in it’s future unless they manage to get ahead of it and change the law, which is exactly what OpenAI recognized and started trying to do.
Hopefully this is obvious to you, but following people on Civitai doesn’t prove anything about how often people try to prompt for specific artists, and given the growing exposure to copyright issues surrounding this, it’s entirely possible that the people using prompts that intend to break copyright are choosing not posting their output. You have absolutely no idea how many people are prompting for specific situations or styles or names. And again, the “majority” usage is completely irrelevant to this discussion.
You've shifted the goalposts so much that I genuinely don't know what you're arguing for anymore. Are you saying "some people use <technology> for <bad>, therefore <technology> is <bad>"? Because then we're back to banning knives, banning cars, banning internet, banning everything basically.
Typically in these discussions it matters what the intended or majority use case for a technology is. If it doesn't matter, then what is it that you're arguing for exactly?
I fully agree with you that some people use image generators to generate new images in specific artists' styles and I agree that there's some moral issues there.
> it matters what the intended or majority use case for a technology is.
Okay, speaking of moving goal posts, you previously made unverifiable claims about what the majority type of prompt is, in order to contradict my points about specific art prompts being common. Until now, our discussion was not about the majority use case for the technology. The prompts and the tech are not the same things. The majority use case of AI technology has nothing to do with whether a user is seeking generic photographs or art in a specific style. The intended use case of Stable Diffusion, for example, is to generate images from a text description. All text prompts and all image outputs fit the stated majority use case of Stable Diffusion. The intent of it has - deliberately - never included any statements on what the prompt contains nor what the output looks like, and half of the examples on Stable Diffusion’s home page are art anyway.
You claim accidental infringement is an unintended failure mode, but it’s a surprise to nobody that a sophisticated copy machine sometimes copies things. Plus they had to copy stuff in order to train it, which may well be illegal even before anyone starts inferencing. It’s fundamentally a copy machine they built, and they intended for it to copy things, so unless they design it to guarantee remixes of many, hoping it won’t copy enough of any single source to violate copyright is wishful thinking and nothing more. The copyright office may decide to view that as willful ignorance in violation of the law. Accidents and intentional infringements both are bound to happen, and the machine is effectively distributing copyrighted work and could be violating existing laws.
So far you’ve ignored all of my points touching on the possibility that training AI on copyrighted material might be illegal. There might be legal issues with AI, unless the people building them curate legal training sets.
You’ve also failed to engage with any points about OpenAI requesting copyright exemption, which I’ve mentioned multiple times and was my first point. Their actions have already demonstrated what I’m saying here, that they know the potential copyright problems with their AI go far beyond overfitting accidents. We haven’t even discussed ChatGPT, but it’s pretty obvious that they would like to have the right to disseminate information and excerpts from copyrighted sources. They want it to overfit on some things, otherwise they can’t improve it’s ability to tell the truth and not hallicinate false information.
You made many points in this thread, and I didn't argue all of them. I don't know why you want me to comment on this specific point? Fine. I agree with you that training AI on copyrighted material _might_ be illegal. In fact, it probably will be declared illegal _somewhere_ (and probably will be declared legal somewhere else). I'm not a lawyer, so I don't have much to contribute to this particular point, not sure why you expect me to comment on it?
> You’ve also failed to engage with any points about OpenAI requesting copyright exemption, which I’ve mentioned multiple times and was my first point. Their actions have already demonstrated what I’m saying here, that they know the potential copyright problems with their AI go far beyond overfitting accidents.
I agree with this as well. I didn't comment on it, because it was tangential to the point that we were arguing about, and I have nothing of substance to contribute to OpenAI's copyright exemption request.
> We haven’t even discussed ChatGPT, but it’s pretty obvious that they would like to have the right to disseminate information and excerpts from copyrighted sources. They want it to overfit on some things, otherwise they can’t improve it’s ability to tell the truth and not hallicinate false information.
That also makes sense.
> You claim accidental infringement is an unintended failure mode, but it’s a surprise to nobody that a sophisticated copy machine sometimes copies things.
Sure. I wouldn't call it a copy machine, but yes, nobody is surprised that the machine _sometimes_ copies things.
> I’ve been suspecting for the entire thread that you’re really not understanding nor listening to my argument
I still don't understand what the "argument" is. You've made several different statements which I agree with. And then you made one outlandishly incorrect statement, which I disagreed with, and that was the point I argued with you. What you're doing right now is muddying the waters with a thousand unrelated arguments, and I don't know why you're doing that or what you're hoping to achieve.
If you can articulate what you believe it is that we disagree on, please go ahead. If you don't know what the disagreement is, then maybe there's no point continuing this conversation further.
Copying the mere style of <famous_artist> has never been a copyright violation.
Ever.
I believe that's answered in the first line of the article.
This just seems like another way to stifle innovation