Artists are fleeing Instagram to keep their work out of Meta's AI
washingtonpost.com
washingtonpost.com
Even if we have strong, even overbearing legislation put into place to protect artists, we're just going to end up with the biggest most profitable offenders buying/using illegally trained models through middle men, and feigning ignorance when/if they are found out.
Though most of the "has this model been trained on X" conundrum is likely to be irrelevant on a years (months?) timescale as art styles are far less unique than artists would like to think. See tencent's PhotoMaker or other modified CLIP approaches. Even a model not trained on a particular face can be conditioned to generate said face using a oneshot approach because the vector representing that face can be constructed even if the face itself doesn't exist in the training data. I'm certain the same is true for artistic styles.
"AI" tech? yes we can, scrape the entire web, fire up the transformers (aka roll that VC money)! "prevent stealing from creators" tech? nah, too difficult, we don't know how, human problem (the neighborhood queer artist can't pay nearly as much).
Fingerprinting technologies. Generative tools can be mandated to embed watermarks. Original images can include fingerprinting and other counter-theft measures that tools must respect. Offenders who publish tools that don't can be prosecuted a la tornado cash. etc
It's not about art style. The model exists thanks to data, period. No data, no vectors. And almost all of that data is stolen.
Additionally, there's no way you will be able to force people to insert fingerprinting into the model or surrounding tooling. There is no way to force fingerprinting into the source material, glaze etc are memes that are easily defeated and in fact can help stabilize outputs when training on glazed images is used as a negative embedding. The idea that any of this stuff is a viable method of combating gen AI is laughable. You can run this stuff on an 8 year old GPU with a blob the size of a couple netflix episodes. While outputs straight from SD may be identifiable, SD + manual intervention can easily create outputs that are not traceable. There is no workable solution that doesn't address the underlying economic incentives.
It doesn't matter that the model is trained on data that you believe is "stolen", people who want the model will have access to it, you won't be able to prove the outputs are AI, and even if you could, it won't be cost effective to prosecute people for creating in this way.
The problem here is our economic system and it's reliance on scarcity and markets to assign value to things that are fundamentally not scarce (AI models, digital goods) and it's failure to correctly assign value to things that have enormous positive externalities (artists creating culture).
You'd be surprised. Pass a law requiring ClosedAI to start doing it tomorrow and what will they do?
If you have the resources to scrape the entire web and sanitize the data and fight the anti-AI arms race and retrain the model all the time and you do it alone just for yourself at home for the lulz then I guess knock yourself out, no one will mind. But a gaming computer from last decade alone probably won't cut it
> entirely open source stack
I heard of some tornado cash and wow using OSS for crime totally made it not a crime!
> millions of other people
Yeah sure my phone also has ML
I can train a LoRA on a particular artist in about 6 hours using my old ass computer and a base model that's almost 2 years old plus 20 images scraped from wherever and achieve great results. Most of what gives the model the generic capability to create images doesn't need to be retrained, finetuning on styles is very cost effective. Though many of the advances in the capability of SD have nothing to do with training the model at all but rather advances in instrumentation and conditioning. Art is like 90% building on generic work that others have done before you and like 10% your own contribution if that. Everything is a remix. That is the the core truth that allows style transfer approaches to work even if the specific style is not represented in the training data, in most cases there's a blend of conditioning vectors that gets you pretty close.
Maybe the future of art is "making art that's hard for the computer to recreate" which would be dope, but the stuff it's already good at? That's over and done with, that cat is not going back in the bag. And to be clear, the set of stuff it's good at is larger than the set of things that are in the training data because of advanced conditioning techniques like CoAdapters/IPAdapter/ControlNet/modified CLIP conditioning etc.
And yes, millions of people, SD 1.5 had 4.5mil downloads in the last month alone.
Promising jail is that point.
Remember, open source is besides the point, whether what you do is against the law is what it's all about;)
> training on an image
Talking about training on a single image like it's enough means you probably don't get the point of ML m8
While I wasn't clear, you're also willfully misunderstanding me. The point is that none of the adversarial tools work, and I have strong doubts that any of them will ever work as a fundamental facet of how the data is processed by these models.
Though you may be surprised that you can make a style LoRA that preforms well off ~8 images and you can condition a model off a single image. Though if you're still doubting me I invite you to pose a challenge for my 8 year old GPU. ;)
Press 'M' to start a Crusade.
your comment also reminds me of discourses by Epictetus. he sounds like an irritated teacher modelling what patience looks like.
she was not the one telling me to learn to be happy with less - I demanded that for myself.
if your problems are financial, you need an accountant. if you struggle to speak to a large crowd, you need a public speaking coach or training group.
a therapist is the one you talk to when you have stuff weighting you down and when you try to talk to others they give advice that starts with "you just need to ...", like you haven't tried it. if reading my post has riled you up, paying somebody to listen to you complaining about it might work quite well.
On the other hand we're starting to see some crazily impressive progress on the physical robotics front.
I think it's definitely still an open question whether we'll see e.g. a humanoid robot that can load a dishwasher, or an LLM that can reliably .. I dunno, edit a video, first.
We have that, it's called a dish washer. We also have a robot for vacuuming, it's called a robovac, and so on.
We have robots for most mundane repetitive chores, they're just not humanoid like in SciFi, but are purposely designed for each task.
But people can also copy you by hand if they want to.
For any other output, even if it does not reproduce 1:1, the only reason it works at all is thanks to acquired training data and most of it is stolen.
https://www.theverge.com/2012/12/20/3790560/instagram-new-te...
Whether there's even any data behind this or if it's an article based on a handful of random tweets who knows, as it's a pay-walled article.
On a meta level, what percentage of HN readers are paying for a subscription to the washington post? Why are links to paid articles even allowed when the vast majority of readers won't even have access to read it?
HN rules explicitly say we aren't allowed to complain about that
Admins seems to be quite open and helpful from what I saw.
I agree with their complaint.
At first I thought it's a scam but after the second artist who was clearly not the type I gave it a look.
I like my little green squares
I can recommend gitea (easiest to setup) or gitolite + cgit (minimal)