Stable Diffusion Is the Most Important AI Art Model Ever
thealgorithmicbridge.substack.com
thealgorithmicbridge.substack.com
Is it just me, or does anyone else think that this is an impossible and futile task? I don't have a solid grasp on what kind of censorship is possible with this technology, but the goal seems to be on par with making sure nobody says anything mean online. People are extremely creative and are going to find the prompts that generate the "harmful" images.
Reminds me of the mysterious control of Conjoiner Drives in Alastair Reynold's books.
The thing about that is that it is open source, so you can trivially disable that filter if you like.
https://github.com/CompVis/stable-diffusion/blob/69ae4b35e0a...
You can turn it off if you want to; it's a simple convenience.
Even in the 90's they had to fight hordes and hordes of Californian nutjobs (Diane Feinstein et. al.) that wanted to ban violent video games. These people would be certainly cancelled in today's world, wouldn't hold a chance. Because, how dare you allow violence in video games to ...children!?
Our civilization depends on allowing wacko's do their thing as far as it is within limits of the law. Let them be offensive as fuck. These are the people that herald and propel society forward by their heterodox thinking. Society is going to decay fast, it already is.
Lots of people from other states, including Texas, in that list too. It wasn't just a California / Left issue.
> During the U.S. Congressional hearing on video game violence, Democratic Party Senator Herb Kohl, working with Senator Joe Lieberman, attempted to illustrate why government regulation of video games was needed by showing clips from 1992's Mortal Kombat and Night Trap (another game featuring digitized actors).
https://en.wikipedia.org/wiki/1993_United_States_Senate_hear...
Could be true, maybe, but today conservatives have willingly taken over that seat, and the NRA is heavily involved and actively blaming video games after each mass shooting to deflect from the debate on gun rights. https://www.usgamer.net/articles/the-nras-long-incoherent-hi...
In terms of trying to moderate swearing and sexuality in games and music and movies, the religious right has long been and still is the group most vocally opposed to such free expression... if we’re talking about where to address censorship today.
I really should have left out the Diane Feinstein and "California nutjobs" in the original post. This is what happens when you mistakenly poke HN every single time when it comes to political one-sidedness.
The fault is mine.
Perhaps a more important discussion, if you do care about censorship, is to define more thoughtfully what you mean about “within the limits of the law”. In the US, the law, up to and including the constitution, makes clear that offensive behavior is anywhere from not protected free speech up to criminal activity. Politicians are debating what the limits of the law should be, and sometimes they blow hot air, and sometimes they write bills. Either way, the results of Congressional bills are establishing the limits of the law, and so define the acceptable legal bounds of offensive media & speech. Here’s one of the bi-partisan congressional sessions on games (it included Feinstein, among many others, but she didn’t testify). https://www.govinfo.gov/content/pkg/CHRG-109shrg28337/html/C...
Only one California politician has ever attempted to do much of anything to video games: republican Joe Baca who tried a dozen times, and is mostly famous for his attempt in 2009 to get a warning sentence on boxes. Calling that censorship is pearl clutching
The only genuine attempt to do something an adult would consider censorship to video games were Jack Thompson, now banned republican, or that brief 2018 thing with Trump.
Democrats have never attempted to censor video games. All three major attempts were Republican.
It's important to get the details right if you are going to build an intuition of who's actually doing this
The evidence is easy to look up, and you didn't give me any.
Anyway, it won't be possible to contain it. Better spend the effort on how to deal with bad actors instead of trying to restrain the use of content creation tools.
The issue with these neural networks isn't the content they create, it's that they can create massive amounts of content, very easily. You can now do things like: write a Facebook crawler which photo-shops people's photos on nudes and sends those to their friends; send out mass phishing emails to old people with pictures of their grand-kids bloody or in hostage situations; send out so many Deepfakes for an important person that nobody can tell whether any of their speeches is legitimate or not. You can also create content even if you have no graphic design skills, and create content impulsively, leading to more gross stuff online.
Spam, misinformation, phishing, and triggering language are already major issues. These models could make it 10x worse.
I literally read a chapter of Inhibitor Phase where there's a ship called "John the Revelator" less than an hour ago. I haven't otherwise seen that phrase written down for years.
Spooky (and cue links to the Baader-Meinhof Wikipedia article).
Or 10x better, as the barriers to entry for doing this kind of thing right now aren't high enough to make it not happen... they are only high enough to make it sufficiently hard to pull off that people can feel comfortable assuming that most of the content they see is legitimate; in a world where nothing is necessarily legitimate I'd expect you'd see a massive shift in peoples' expectations.
I immediately came up with "Call the football team, I'm wet" and "Daddy lets play hide the sausage" as example workarounds.
It's entirely pointless. Humans are vastly superior in their ability to subvert and corrupt. Even if you were able to catch regular "harmful" images humans would create a new categories of imagery which people would experience as "harmful", employ allusions, illusions, proxies, irony etc. It's endless.
Complicating matters more is the fact that something being censored can be considered harmful as well. Religious messages would be a good example of this; Religion A thinks that Religion B is harmful, and vice-versa. I doubt any 'neutral network' can resolve that problem without the decision itself being harmful to some subset of people.
While I love the developments in machine learning/neural networks/etc. right now, I think it's a bit early to put that much faith in them (to the point where we think they can solve such a problem like "ban all the harmful things").
There's way too much moralizing from people who have no idea what's going on
All the filter actually is is an object recognizer trained on genital images, and it can be turned off
The prompt isn't very relevant. I've had the filter fire on completely innocent text
The filter is a simple checkbox in preferences. All this deep thought is missing the point. You can just turn it off
The model was trained on eight specific body parts. If it doesn't see those, it doesn't fire. That's 100% of the job.
I see that you've managed to name things that you think aren't in the model. That's nice. That's not related to what this company did, though.
You seem to be confusing how you think a system like this might work with what this company clearly explained as what they did. This isn't hypothetical. You can just go to their webpage and look.
The NSFW filter on Stable Diffusion is simply an image body part recognizer run against the generated image. It has nothing to do with the prompt text at all.
It is obvious to anyone who bothers to try - have you? - that a filter was placed here at the training level. Rare activities such as "Kitesurfing" produces flawless, accurate pictures, whereas anything sexual or remotely lewd ("peeing") doesn't. This is a conscious decision by whoever produced this model.
>All the filter actually is is an object recognizer trained on genital images, and it can be turned off
I'm not sure if you misread something, but neither I or the person I was replying to was talking about this specific implementation, but in a more general sense?
I'm pretty sure you are the one who missed the point of the parent post and mine.
Anyway, the filters you're thinking about don't exist and you can download the code and use it today.
Thanks for speculating.
Who is "we"? Do "we" consider nude paintings to be harmful?
Is "we" Mike Pence? Roman Polanski? Woody Allen?
There is no coherent "we" and no consensus on what "we" consider harmful, so no AI can possibly learn that.
Isn't this part of the AI alignment problem? To be able to understand what kinds of output is unacceptable for a certain audience? To be polite?
Do we want the AI to generate based on Polanski's sensibilities, even if he's the only audience member? I suspect for most people the answer is no.
Automatic review of content, NSFW filters, SPAM filters etc... have been bog standard since the earliest days of the internet.
I don't think anyone likes it. Some fight it and create their own spaces that allow certain types of content. Most people accept it though and move on with their lives
SPAM filters on SMTP ports were implemented long long before any government mandated it - at the ISP level often
Further, the development DKIM and SPF, were incidental to any government requirement
Preemptively: The fact that the early internet was heavily government and institutionally focused, doesn't a government mandate make
So if the corporation is an intelligent collective, then it's regularly outsmarted by other intelligent collectives determined to bypass it.
e.g. the one line of code in Stable Diffusion that predicts if stuff is NSFW, can be inverted to generate only NSFW stuff.
I tend to agree with OP that there is no technical solution to this problem.
Problem with tech is once it’s known to be possible if you choose to try and monetize it by making it public as OpenAI and Google were planning to do then it’s only a matter of time before another smart team figure out how you’re doing it.
You can do the Manhattan Project in secret and in 500 years someone else might not realize it’s possible. But the second you do a test of that concept the sign you did that is detectable everywhere and the dots of what you did will connect in someone’s brain somewhere.
Can’t put the genie back in the bottle.
This is employing a fallacy that people have infinite amounts of energy and motivation to devote to being hateful. I have been on countless online communities in video games and elsewhere and when the chat in them doesnt allow you to say toxic, hateful stuff... guess what a whole lot less of that shit is said. Are there people who get around it by changing out characters to ones that look the same that dont trigger the censor or by using slang or by mispelling? Of course but the fact is I think if you talk to someone who runs communities like this they would laugh in your face if you said a degree of censorship of hate speech wasn't fundamentally beneficial.
A big aspect has got to do with the fact that if everybody agrees to be part of a community, part of that agreement is a social contract not to use hate speech and if someone flaunts that they are bypassing it.. in the obvious flaunting of the social contract established (it is obvious they had to purposely mispell the word) these people are alienating themselves by underlining the fact that the 99% of the community finds their behavior pathetic and unacceptable.
I think spaces that effectively moderate AI art content will be successful (or not) based on these same factors.
It won't depend on some brittle technology for predicting if something is harmful or NSFW. (Which, incidentally, people will use to optimize/find NSFW content specifically, as they already do with Stable Diffusion).
However the point stands that as a concept, humans will find a way to exploit and corrupt any technology. This is unquestionably true.
Bertrand Russell famously makes exactly this point as well, albeit specifically when it comes to violent application of technology in war. That: until all war is illegal every technological development will be used for War.
Your point however is also true, in that in certain spaces for certain audiences (communities), participants make it more difficult to exploit these things in ways that they don't want to and to explout them in ways they do.
Ergo, Technology is and remains neutral (as it has no will of it's own) and the people using and implementing technology are very much not neutral and imbue the will of the user onto the tool.
The real question you should be asking is, how powerful can a free tool/knowledge get before people start saying that only certain class of "clerics" can use it or that most communities agree that NO community should have it.
Notice on that last point how not-hard we're trying to get rid of Nuclear Weapons
If I swear at a video game and it comes out as ** I might think "OK, maybe I'm being a bit of an asshole, there could be kids here and it's a community with rules so I'll rather not say that".
If a tool to make art doesn't let me generate a nude because some American prude decided that I shouldn't, though... my reaction is going to be to fight the restriction in whatever way I'm able.
What a great insight which I find both extremely poignant and simultaneously disheartening
Perhaps this will come as a comfort to the people who are vehemently against creating human-level AI systems
Useless effort.
They spent years grappling with online worlds because of the idea that people might/could represent themselves as a different gender, they wanted the technology to exist and had dreamed about it for decades they just got caught up on that
That was comical because it was also out of touch at the time period as well
Its interesting how people squirrel and spiral over useless things for some time
Humanity has had the ability to lie with pictures since the invention of photography. The field of special effects can be described as lying about things that don't matter.
Without using Stable Diffusion, I can still photoshop an image or deepfake a video. Stable Diffusion isn't really changing what's possible here, and arguably is less advanced than what's possible with Deepfakes or even the facial filters available on social networks.
Like with all deceptive imagery: one just needs to use their noggin.
* Also I might add: the article is actually out of date on some aspects, because this technology is evolving so rapidly. Literally every day there is a new and interesting way that people are applying the tech.
In both tools you can get naughty images, but you have to tell the tool that's okay.
This is not about censorship or moralizing.
It is just having the tool know when it's allowed to do that stuff. It's a key basic product feature if you're actually using the thing for content and not just having fun making pictures
Everyone acting like there's some kind of free speech issue should go into their account and turn the filter off, then try to calm down
If there’s anything we’ve learned from history, it’s that we’ve always been morally wrong in some way, very often in our most strongly held beliefs. This AI in a different time would be strictly guided to produce pro-(Catholic Church/eugenics/slavery/racist/nationalist) content.
It’s not just morality - there reportedly have already been multiple subreddits of non-consensual porn trying to mimic real people and underage porn. The legality of that is a minefield but it doesn’t end there. If that’s what they become known for it affects funding, hiring, people deciding whether to use their software, etc. and the more prominent that is the more likely that they’ll be hauled before legislators to talk about problems. Even simple things like legal demands to remove celebrities from the training sets could be pretty time-consuming.
Then a new inmate joins, doesn't know what's going on but figures that if you say a number, people laugh. So he goes "14!". But nothing happens. The others tell him "you didn't tell the joke right".
How is the poor AI meant to know that jokes 6, 13 and 38 are sexist?
“It’s all in the delivery.”
That being said, I'd still think publishing the model (vs keeping it as a closed-source API) is a good move. Otherwise, we'd move forward into a world where one of tge most significant technological advancements must be gatekept forever, which I'd frankly find even more dystopic.
The goal is to have a checkbox which keeps the system from generating naughty images in casual use.
This has absolutely nothing to do with censorship. It's a nonsense concept and it's not clear what you think censorship actually is.
If you set the system to make tall rectangles, are you censoring squares?
It's absolutely exhausting how people on HN attempt to cast any form of telling a tool what you want the tool to make as if you're somehow morally governing something
It's just telling the machine what to make
Not everything is a desperate ethical dilemma
Sometimes you just want the things you create to be straightforwardly usable
You understand that the filter is voluntary, and that the initial delay requirement (long gone) was about Discord adult image rules, right?
You're not just moralizing censorship by habit where there was none, trusting hn to overreact when that word was abused, right?
If we can't even do this, how are we ever going to align AGI? I see these efforts as part of a nascent effort at alignment research (along with the more proximate reason, which is avoiding bad PR from model misuse).
Like,
View: 50mm film, wide-angle
Scene: rectangular room with window -> show preview
Scene: add table -> show preview
Scene: move table left -> show preview
Scene: add mug on table -> show preview
View: center on mug
Right now, there’s little control and it’s a lot of random guessing, “Hmm what happens if I add these two terms?”
For example: https://www.reddit.com/r/StableDiffusion/comments/wwgge8/ano...
Consider also this example of someone splicing Stable Diffusion into a proper image editor and using a combination of img2img, text to image, inpainting, and normal photoshop tools: https://www.reddit.com/r/StableDiffusion/comments/wyduk1/sho...
What would be best is to properly integrate models like that into some painting software like Krita. Imagine a brush that only affects freckles, blue teapots, fingers, or sharp corners. (or any other thing in a prompt) Or a brush that learns your personal style and transfers it onto a rough sketch you make, speeding up the process. Many possibilities.
I think they are already making an img2img plugin for Photoshop. Watch the demo, it's kind of impressive. [0] It's just a rudimentary prototype of what's possible with a properly trained model, but it already looks like a drop-in replacement for photobashing (as an example).
https://old.reddit.com/r/StableDiffusion/comments/wyduk1/sho...
Turns out the Star Trek predicted 2020's style AI behaviour rather well. Considering nuclear war is then due in 2026, that's disconcerting.
I find the best holodeck “prompt” scene to be Picard explaining how he’d like to experience the world of Dixon Hill in “Manhunt”:
Still seems odd that it's only apparently Data, Moriarty and the Doctor that have demonstrated the Federation actually can make pretty general AI with the tools it already has on starships (and conveniently always on the ship with all those film crews on it making the Historical Records).
Surely under the crust of some demon class planet there's a bank of millions of times that power bring used for...something.
There's probably a rule against making AI that you're allowed to break in the delta quadrant though.
And there is a direct canon line from Moriarty through to the EMH and later sentient holograms via Lt. Barclay.
FTFY
Dall-E only really became known after Dalle-E Mini was used to flood the internet with memes.
If you want a set of images with the same artistic style for example, especially a distinctive one, that can be hard to do
If you want a set of images starring the same recognizable character or object, in eg. different situations, that's gonna be real hard to prompt
More control is coming: https://dreambooth.github.io/
What I am actually leaning towards is a tool for users to "enhance" art with AI, but only if the artist allows it.
Currently it takes my mid end PC (2070 Super) 10 seconds per image, which is too slow. You would need to get generation time below 1 second to be quite productive. I guess you can already achieve that with something like triple 3090s?
https://www.reddit.com/r/StableDiffusion/comments/wyduk1/sho...
Doesn’t match the workflow you’re describing exactly, but shows how this stuff can be integrated somewhat smoothly into a UI.
If the point of all of these models is to get to something resembling an artist, then why intentionally kneecap it from the start and prevent it from producing art?
I'm not sure, but they mention biases. So I imagine one thing they want to avoid is that you ask for a drawing of a "criminal" and 90% of the images are of colored people. It should be possible to avoid to minimize these if you review the dataset, at least for certain key words.
Also, you can fine tune it on whatever you want which is awesome.
One interesting effect I have noticed on myself though is that after staring at DALLE or Stable Diffusion generated images for a long time then viewing "real" media, I get the same sense of wrongness that the output is not quite right for awhile, like my brain has been tweaking its processing to prefer the AI art as the ground truth!
Also, there's no content filtering, but I don't recommend playing around with that if you're sensitive. The lifeless husks and various mixes of body parts I got when playing around with it with fairly benign phrases could very well be used for a horror movie.
It might be that I haven't yet found the right phrase for stable diffusion for pencil sketches though, as for dalle2 it's just "<describe what you want>, artstation, pencil sketch, 4k" to generate consistently great pictures.
https://www.instagram.com/p/Ch0PHzvvXgr/?igshid=YmMyMTA2M2Y=
https://www.instagram.com/p/ChwuZlmuJSV/?igshid=YmMyMTA2M2Y=
https://www.instagram.com/p/ChwZmUhukK_/?igshid=YmMyMTA2M2Y=
https://www.instagram.com/p/ChusQWELB-x/?igshid=YmMyMTA2M2Y=
https://www.instagram.com/p/ChureAOL2H0/?igshid=YmMyMTA2M2Y=
The first links are most recent. You can see the progress I've been making as I learn to do better prompt engineering and iterate on existing images by using img2img. The future is here...
> To generate or disseminate verifiably false information and/or content with the purpose of harming others;
How do they define that? So I can generate disseminate false information without the purpose of harming others, just for fun? And what if I believe I am not harming others, but helping others? Can I generate fakes to further my political cause, if I'm convinced it is a "good" cause? And what about Popper's paradox? If I prevent people from harming others, then I am still causing the harmer's harm. I feel they are opening a can of worms here.
Also, bad actors will just ignore the license. There is a piece of code that censors obscene generations, you could just comment that out. I feel the license and that filter are not going to stop anybody, but are mostly there for good publicity and so they can wash their hands in innocense...
at the bottom another linked article:
> BLOOM Is the Most Important AI Model of the Decade
All very interesting but it's a bit hard to take serious if they exaggerate the titles so much that they contradict each other
> "You may be wondering if such a bold headline is true. The answer is yes."
Feel pretty save to tag 'The Algorithmic Bridge' as not something to pay attention to in the future.
Honestly if I was a commercial concept artist or illustrator that didn’t have a signature style I’d be really worried. We’re truly gonna see the power of this tech as a tool now it’s not gatekept.
The prices people pay for any kind of picture art are about to take a nosedive. Stock art websites are going to be hit hard, any kind of graphics artist, any kind of commissioned artist. I wonder if (human) models will be taking a pay cut as a result.
While it's a huge win to be open source, I find the results always inferior to Midjourney (and DALL-E).
I tried to generate some artistic results with variety of prompts and Midjourney always won hands down.
But of course, since it's open source, many community tweaks and colab notebooks/forks will probably put it in par with DALL-E by time. But I have trouble imagining Stable Diffusion competing against Midjourney anytime soon: the different is day and night.
It was so bad that if they'd replaced v3 with it (good that they didn't) I'd probably have stopped using MJ and had cancelled my subscription.
Even the team acknowledges that beta model isn't for everything and v3 isn't going anywhere soon.
As I argued there, AI art models like Stable Diffusion pertain to a new category of tools and should be understood with new frameworks of thought adapted to the new realities we’re living in. We can’t simply make analogies or parallelisms with other epochs and expect to be able to explain or predict what it’s going to happen accurately. Some things will be similar and others won’t. We have to treat this impending future as uncharted territory.
I wonder if we'll also talk about "conversations", "complex situations" and "the need to treat this as uncharted territory" when some Copilot/GPT3 successor a few years down the line spits out entire production-ready software stacks off the prompt "like Facebook only better" - using our own code as training data.
This prompt is unspecific to the point of unusableness. Even if this works some day, the spec used will be a lot more detailed, in higher-level pseudocode style.
There will absolutely be prompt engineering and I agree that actual, serious prompts will be much more specific than that.
I don't think the prompts will necessarily be pseudocode-style. Depending on what trainsets are available, I could imagine we'll have some high-level description of desired features in addition to lots of specifiers which narrow down the specific languages, design patterns, tools etc which should be used in the resulting codebase.
You can already use similar prompts with Copilot today by disguising them as comments.
I'm just a run of the mill software engineer (mostly webdev).
I never cared about ML, or Data Science,
Had been playing with dall-e the past few weeks after getting beta access, but it's too limited/meh and ran out of credits soon.
Then DreamStudio (SD SaaS) launched to the public, and I was blown away.
Then I tried to run txt2img on my mac, which I did, but it's too cumbersome/slow
Then found out about replicate, which also exposes an API to interact with and run the models.
I've been since then having fun with it, built some scripts with playwright, and doing generative art with stable diffusion, I'm no artist, but it's so much fun, and results so visually pleasing, I cannot not pursue the urge to explore this.
I will be starting an anon account in twitter and try to sell some of my art as NFT's, we'll see where it gets me.
Just ordered a card to get around NSFW filters (They' are nonsense and flag some random stuff)
If you want to try it the easiest way is dreamstudio/replicate.com
… attempting to limit the generation space to omit porn, copyright infringement, violence, racially “unbalanced” content … etc.
I think looking at text prompts as an essential part of the technology is very limiting. You could be using a mix of text, storyboards, images of objects to place inside it, sketches of the desired layout, etc… in fact why not replace your game's renderer with it?
What astrange idea, surely it would never work? /winkwink
1. Sharing is publicly is still considered highly illegal. 2. The biggest problem with everything CP related is that to produce this kind of content, real children were used, tortured and often killed.
Should we not allowed the model to generate CP content in order to reduce the number of real children being hurt/killed?
Plus, this of course would not prevent any authorities from tracking who is sharing this kind of content.
These licenses will do little to nothing to stop abuse: The abusers will already conceal their identities because their actions are immoral or even illegal (fraud, harassment, etc). But they create a whole host of new liabilities for the users because the definitions are exceedingly subjective.
It's tremendously important to make these tools actually open. But open with a lurking liability bomb stops short of the goal. While stability.ai may never turn into a troll or sell their rights to one, that isn't necessarily true for the next model that comes around.
I guess it matters because most of these images are unusable for further purposes because they can't really be edited and touched up easily to fix up all the flaws or do the final adaptation. Are there any options that generate the images in anything like vector art that would then facilitate the downstream finishing process rather than fully rasterised bitmaps?
I feel this would help a lot with giving artists more credit with the AI art outputs.
My only gripe is the usual one in the AI field: very sloppy nomenclature.
Reading about diffusion models I first expected a novel parametrized family of functions, otherwise known as an "architecture".
Instead it seems more like a training method, so a nomenclature of "diffusion training" would seem more apt.
(Direct link doesn't work, count yourself lucky)
Still fascinating how these DALE-E and Stable Diffusion models work though