Nick Clegg says asking artists for use permission would 'kill' AI industry
theverge.com
theverge.com
> "And by the way if you did it in Britain and no one else did it, you would basically kill the AI industry in this country overnight."
This is important to keep in mind. There are competing AI R&D orgs in dozens of different countries, not just America, China and France.
This "logic" that you need to do obviously illegal stuff (under current legislation) because otherwise someone else would do it and take the loot from that raid remains highly questionable no matter what country you put in that placeholder.
The same "logic" would for example dictate that you should dismiss patents as otherwise "someone in China" could do it before you. Somehow nobody would agree that this is a valid approach in that case…
If you’re the leader of a nation, that’s the reality you need to work within. I really don’t care, as someone responsible for 330 million souls, if everyone is pretty equally robbed to support the whole nation. That’s like… taxes, man.
The latter is something we have no problem with in our culture. Every generation of artists do this. Every artist does this, in fact. It may be how creativity itself works, for all I know. But we never call it stealing, or when we do we say things like "great artists steal" and acknowledge it as a process rather than a problem.
It feels to me like when people make the "AI is stealing" argument, they're really arguing about something else and using this as a proxy for it. Not exactly sure what that other thing is. Fear of losing work and being displaced is the obvious one, and it's a valid fear. I share it! But, if that's what it is, I think we should be having that discussion directly, rather than talking around it. We always do the thing where we piss around talking about proxies, rather than going right to the agon.
“I’m original because I add the little curls to the lines, make the lines softer, and make faces pointy rather than round” isn’t originality, nor is any of the art that people claim copyright for.
Additionally, it is a fundamental misunderstanding of AI modes if you think the models are simply lossy compression of data. That is like saying the alphabet is just lossy compression of The Odyssey.
One could of course say that "IP" ("intellectual property") is an ill concept all in all and abolish it. But than without exceptions! If it's not a valid concept in regard to "AI" it's not a valid concept in general. But if we want to keep that concept "AI" needs to bend to the same rules that apply to anybody else. Simple as that. Law is universal. (At least it should be like that everybody is equal before the law.)
What we’re taking about is suing the tool that could potentially be used to generate copyright violations. Which seems silly?
Steelmanning your argument a bit more: OpenAI is selling the fact that you could potentially generate copyrightable material as a subscription, because the artwork is somewhere in the compressed data?
Abstracting the underlying tech a bit: If we imagined a printer that could faithfully represent whatever you told it to and you told it to reproduce Da Vinci’s that you later sold as originals while others told it to reproduce original forms of art (successfully), who’s at fault here, you or the printer?
Seems obvious to me it’s the user of the printer, not the printer itself, even if the printermaker is making money.
Now you might say that's an overly rigid reading of copyright law, and I might agree, but that is nevertheless the reading of it applied to other domains - music, film and so on - where incorporating even a fractional, lossy and heavily processed portion of another work is enough to trigger copyright law. There is no blanket de-minimis that says "well, it's only 5%, so it's fair use".
Though most of these cases aren't litigating this. They are charging that the ability to recreate copyrightable material is the fault of the tool and not the tool user[1].
If I measure the length of a song? Not copyright infringement. If I take even a very short sequence of PCM samples, copyright infringement.
But the practical thresholds appear to be different, there's a bunch of ways you can extract kbs' worth of information from a song, some are considered copying (ripping the melody to midi with a converter, say), some are not.
And in the case of "AI", depending on the model it could be extracting and encoding almost anything, because it's a catch-all term encompassing dozens of architectures and methods.
A drop of blood falls in a lake.. which bit of lake now contains blood?
The user is in this case the "AI" company.
They already violated law almost certainly by copying stuff from the net in the first place. (Almost certainly as it by now does not look like this was fair use, and there is no other means to make legal what they did.)
Some people (like the NYT¹) are in fact demanding before curt the destruction of the trained models as they obviously contain copyrighted content.
¹ https://pressgazette.co.uk/media_law/new-york-times-open-ai-...
Yeah, that's the express and sole purpose of the tool. You use OpenAI image generation because you want something that looks like a stock image for your website. But you don't want to pay for a stock image because you're cheap. So you'll get a close enough approximation for free. Or, uh, less.
Categorically untrue. 100% of all my usage of image generation is for artistic purposes that have nothing to do with the reproduction of or even the mimicry of existing art.
This is a gaping hole in the legal case if the lawyers need this to be true.
How? You're creating artwork that is based off of the training data and is meant to replicate a subset of it.
I think, just because you cannot pinpoint the exact art piece(s) that you are recreating doesn't mean that that isn't what you're doing.
"just because you cannot pinpoint the exact art piece(s)"
Precisely, and neither can the owner of the copyright. Unless of course I rub into classical copyright issues ("these two art pieces look too similar") which goes back to my original question, why isn't just enforcing existing copyright law on bad ChatGPT users enough?
If I "use" your artwork, and generate another artwork that looks nothing like yours (either from my brain or by automating it) why do you care?
I wouldn't, but presumably the person who you are immitating would.
The fallacy here is that you think an LLM is taking one or two art pieces and then turning around and creating something brand new, like a human brain might do.
No. It's taking billions of art pieces and then sampling them at various frequencies to create your art work. We can find your art work somewhere in the training data, it's just hard.
In essence, we're smuggling the copyright violation through convolution. Since these models are truly very large, we can pretend that there isn't copyright violation. But there is, we just can't really find it because LLMs cannot work backwards.
I don't think the human brain is doing this at all (except maybe if you're writing a satire or homage). Your writing is the sum total of all your experiences which includes everything you've ever read, a number that might (depending on your age) at least approach millions of tokens.
And just like an LLM you'd find it difficult to piece together your influences.
Ultimately these are tools, slaves to humans, meant for one thing only - to serve us. If they're not serving us, they should be eradicated. Because that's what we do.
Stealing from humans, no matter how convoluted, is bad. We shouldn't do that, and luckily that's completely optional and we don't have to stand for it. The computer programs can't fight back, it's okay.
So getting to the bottom of what is actually stealing or not will be important (we can't just unilaterally declare it so), and AFAICT if a user misuses it to regenerate copyrighted material that particular user is at fault, not the tool itself.
AI companies use work that they do not have the rights for, to construct their products.
It's closer to archival and compression than captioning. But they're selling it. That's a big problem - you can't just take someone else's work, sprinkle some god knows what on it, throw it in a big pot with other shit you stole, and then reproduce it later. Even if your reproduction is super convoluted and not 100% accurate.
It's fine to nail a $10m verdict against a guy with no money, but dare to tread on the beloved and bespoke * aI iNdUsTrY * and suddenly you've got better things to do?
That is not a bad thing.
If I use photoshop to create a straight reproduction of an artwork by someone else and monetize it, copyright law handles this.
If I use AI to do the same, same deal? At no point did someone suggest suing Adobe for building photoshop for "enabling" people to recreate art and sell it.
So what changed? The ease of use?
Neither does OpenAI? What about an empty chat prompt “encourages” copyright abuse?
When you ask for a picture of a cup of wine, you're reproducing an image of a cup of wine that was using in training. Or, more accurately, hundreds of images sort of averaged together.
You can't say this isn't the intention because this is literally the singular use of LLM. They don't do anything else. They don't change your oil. They reproduce input. Just because you don't know exactly what the input was originally (and you, in fact, cannot find out) does not mean that isn't what it's doing.
I'd be more interested in getting AI companies training with publicly scraped datasets to have them available to other AI trainers for free. The load scrapers put on the web is out of control.
I actually agree with him on this though, AI is a bigger net benefit to humanity than some artists' IP.
The fact that you refer to an actual human being as “some artist” betrays the absolute disdain you hold for humanity and in my opinion should disqualify you from polite discourse of any kind.
I'm going to make an equally supported and valid statement here and say AI is clearly a net positive to humanity in its current form, and almost certainly all future forms.
You should consider providing a less hyperbolic point to argue if you are looking to have any kind of productive or healthy discussion on the internet.
This is not a productive conversation.
There I fixed it for you. If it's such a net benefit to humanity governments should provide 6-months-behind-the-state-of-the-art-model for free to the public to make up for the highway robbery of its citizens' lifelong digital output.
AI companies are for profit organizations in a largely capitalist world. If what they're providing is so massively more valuable to humanity than artists' IP, then the AI companies should easily be able to monetize that value in a way that allows them to fairly compensate the artists for the crucial role they're apparently playing in AI development.
Of course scale is the real problem here, because we're not just talking about AI companies compensating "some artists", but likely "all artists". This probably is financially untenable since the current value AI is providing to society seems unlikely to be greater than the value of all IP that exists in the world.
But that's not some critical feature of AI, it just speaks to the way current AI models work where you take every bit of text and images that exists and throw it into a data center full of GPUs to burn through a large city's worth of electricity until something approximating human text and image generation pops out. Arguably that approach is driven by the fact that GPUs and electricity are affordable and the input data is "free" if you're willing to steal it. I see no reason to believe less wasteful models couldn't produce good results without requiring the input being all human knowledge, and you'd save some electricity and resources in the meantime.
Edit: I'm not an AI researcher, but I also think incentivizing more efficient approaches to AI model development and training might yield more human-like AI (i.e. actual AI rather than turbo charged autocomplete). Humans also develop our cognitive abilities by absorbing external information but we manage to do it by absorbing a fraction of the information as a large-scale AI model with significantly better results. Focusing on whatever difference is behind that rather than building ever larger data centers with ever more powerful GPUs might be an interesting shift.
All of which to say if Nick Clegg says X there are a very large number of people who think “not X” by default. He’s even a lame duck at meta because they’re ditching him to replace him with Joel Kaplan who they think will help in their attempts to cosy up further with the Trump administration.
[1] Tuition fees for tertiary education in particular https://blogs.lse.ac.uk/politicsandpolicy/libdems-tuition-fe...
This is just Mutually Assured Destruction all over again. The only correct solution is for AI pacifists on both sides (means the people who support humanity and oppose this nonsense) to do a as much as they can sideline the AI hardliners (the people who think AI is a benefit) on both sides. Each country needs to believe the AI pacifists in the other country have relatively more support than the hardliners.
If it was the same for everyone I would have much less problem with this but it's the glaring disparity in how different parties are expected to behave that is truly repugnant.
Working for yourself, and selling what you produce yourself, is called Cottage Industry and is what Capitalism replaced.
It seems people don't like the reality of where we currently are . Don't shoot the messenger!
Licensing law is the opposite: you are not permitted to do anything with a protected work unless you have been granted a prior written license.
(exceptions: fair use for scientific, non-commercial research, and copyright expiry e.g. seven decades after the death of the author, depending on the jurisdiction, of course.)