Stable Diffusion PR optimizes VRAM, generate 576x1280 images with 6 GB VRAM
github.com
github.com
It looks like the committer made their changes in the top commit, then merged the updated CompViz StableDiffusion change set on top of it for some reason. That's where the license change, rick astley image, etc come from.
And yes, StableDiffusion from the original repo will rick roll you if you try to generate something that triggers its NSFW filter.
Here's the code that does it:
https://github.com/CompVis/stable-diffusion/blob/main/script...
And here's what it looks like:
It goes without saying that the authors of a piece of software have the right to make the software do whatever they want, but that shouldn't stop us from recognizing that AI engineers are starting to act like megalomaniac overseers who consider it part of their mission to steer humanity onto the "right" path.
Who exactly do these people think they are?
Imagine this behavior from a web browser. "The URL of the file you were trying to download triggered my NSFW classifier, so I'm going to replace the file with this funny image."
This isn't funny, it's creepy.
Also, if the author wanted to block NSFW contents at all, I'm pretty sure one can actually make the filter inseparable from the main network. This isn't the case here AFAIK.
Don't let people who are clearly motivated by a moralistic desire to control others hide behind BS pseudo-legal excuses.
2. Photoshop's sole function is to let people draw, so drawing a porn w/ Photoshop is 100% users' responsibility. In case of image generators, the AI itself has the capabilities to create NSFW contents. If you're talking about the interactive nature of "prompt", that logic also makes any adult games SFW - they are safe until the user clicks something.
Yes, for malware. As in, software created by criminals to damage your computer and/or steal your personal information.
That's not even remotely comparable to what StableDiffusion does here. What they do is refuse to generate content you requested, based on opaque criteria that no doubt are ultimately influenced by quasi-religious sentiments retained from the Bronze Age.
Why is that a problem, exactly?
> That can cause all sorts of lasting psychological harm and other societal negative consequences.
Like what? To me this sounds like a problem with how some people want to have absolute control over what others do and want everyone else to play along. That's what I find really creepy.
Someone wants to generate porn of me fucking a donkey? Let them. They already could do that with ms-paint and I don't get magically harmed if they do.
This reminds me of the old saying: "Sticks and stones may break my bones, but AI-generated/photoshopped/glued-together pictures of me fucking a donkey will never hurt me" :)
Denial does not make healthy people.
But I do believe that right now we're kinda off track. We almost venerate the act of being hurt. Everyone likes attention and nothing gets such protection by certain classes as having been offended or wronged by some other class. Social signals are currently built to display virtue and so people will go out of their way to display their support of the wronged. I _do_ believe that this is the correct direction to move from where we were, but I think it's gone a little too far and needs time to rebound.
Being a victim is the fastest way to go from zero to hero (reach millions of people) these days and it's also seemingly the least likely way to backfire. People are much more hesitant to bring up the wrongs of someone who's currently being defended for fear of ending up being placed in the out group and ostracized from the signaling group.
Mental health awareness is good. But social signals lead individuals to believing they must be hurt to be a part of the in group. Virtue through suffering is an incredibly effective signaling mechanism
Part of living is opening yourself up to people. That includes the risk of getting hurt. That's normal and part of the human experience.
https://www.theguardian.com/lifeandstyle/2018/may/07/many-re...
The reason why revenge porn, rumours, and defamation are all possible online is because there's a level of plausibility to these things -- especially revenge porn.
Now that plausibility goes out the window because of AI generated content. Someone posts your actual nudes? Give a nervous laugh and say that AI made them.
Someone says they heard a rumour about you and asks if it's true? Say 'don't believe everything you read on the internet, you know they have AI that writes anything you want now, right?"
I'm glad this technology is finally out in the open where people like you no longer have a say in how it is used. The sooner people accept its existence (like they have with Photoshop), the better off and healthier we'll be as a society.
Do you really need someone to spell out to you why being able to create realistic-looking porn of anyone might lead to some issues...?
>Someone wants to generate porn of me fucking a donkey? Let them. They already could do that with ms-paint and I don't get magically harmed if they do.
What an absolutely wild take. Do you expect this line of thinking to be convincing when you're comparing AI-generated images to what someone can whip together in MS Paint? Come on. You, me, and everyone else reading here knows that's not even close to a valid comparison.
A less creative next generation as they don't have to close there eyes and imagine generated porn of anyone.
The second problem will be finding people that you would even want generated porn for.
Yes. I honestly have no
> What an absolutely wild take. Do you expect this line of thinking to be convincing when you're comparing AI-generated images to what someone can whip together in MS Paint? Come on. You, me, and everyone else reading here knows that's not even close to a valid comparison.
Wild or not, my point stands. With minimal skills, you can photoshop anyone's face on any pornographic image out there. It's a spot on comparison.
I would argue the authors went out of their way to actually make it very easy to “decensor” their model in a way that lets them wash their hands of things. So regardless of the outrage over their Puritanism, they actually have made it easier to make pornography than ever before. I don’t think applying the tiny speedbump of their model being censored by default is unreasonable in the interests of conservativism.
In the long run these men will go down in the history of pornography.
As someone that manages nsfw open source projects, this move seems fine to me.
And actually kinda hilarious.
As currently implemented (and the implementation took more work than a confirmation prompt would have!), it's an obvious attempt to control, rather than protect, users. They try their best to dress it up like a joke, but it clearly isn't.
It's not fundamental to SD in any way, and their suggestion wouldn't have worked.
I should add that it’s also incredibly easy to comment it out.
This is quite a classic HN kind of comment. Immediately assumes specific problematic intent and proposes a solution that doesn't fit the API.
They put a simple check in, that tries to avoid returning nsfw images so that you don't get that back despite a more 'innocent' prompt. It's trivial to remove and is only part of the demonstration scripts, it is not part of the model or anything fundamental to it's workings.
> AI engineers starting to act like megalomaniac overseers who consider it part of their mission to steer humanity onto the "right" path.
You outdid your comment. You have an idea of how this should work. And you're trying to supposedly "steer" in a path of humanity. It's just projection.
This kind of preachery righteousness control tactics should sit elsewhere (fork it). The devs want to convey something is nsfw. They will do it however they want.
It's silly and you're able to patch it out if you so desire. But most importantly, they released the model to everyone so it's way more open than OpenAI's DALL-E 2 (for better or worse). I don't think the argument that they're trying to control you really makes a lot of sense given how widely accessible this model is.
It's also clearly in place because they don't have an age disclaimer, which would likely open them up to some sort of liability (or additional liability). Again, it's open source, and freely available, so like. . . this feels like a you problem, sorry.
These people wrote a bunch of Python code that pipes images and text through a GPU, and they're acting as if they had created a secret weapon that somehow humanity must be protected from. If that's not megalomania, I don't know what is.
Where "bad actors" are defined as "people who disagree with us", and "bad content" is defined as "things we don't want to see".
Needless to say, the list of bad actors never includes the authors themselves, and the list of unacceptable applications never includes anything the authors had in mind.
Honestly it just feels weird. I've read the license restrictions and I can't see why they are there out what they're preventing.
(I'm trying to understand the exact difference here.) So it boils down to democratization? That anyone can do it regardless of skills acquired?
Take it out of drawing. If you write a program to control elevators and it breaks the elevators aren't you, the person that wrote the program, responsible? Why would Adobe be responsible for something someone else draws?
Or let's take an easier more close example. If you just made a character generator. Here's one
https://www.numuki.com/game/mii-creator/
And there was a "randomize" button that one out of 100% made a very pornographic image. Who would get the blame? The person that pushed the button or the person that created the project?
1) figuring out how to rejig the prompt to get what you'd like, adjusting seeds and tuning configuration options. The more control you want, the more complex and manual the pipeline of backing software will be. All it does is amplify everyone's innate artistic talent.
2) Coming up with a good prompt. This relies on the person's imagination, facility with words and familiarity of limitations of the image generating software and its training datasets.
3) Selection. This can be a frustrating experience that tries one's patience.
> made a very pornographic image. Who would get the blame?
You would, it's not like the software automatically distributes all its generations. The vast majority of images these software generate are not good enough to be shared and aren't. You made the conscious decision to share it.
Even if it were an AGI, you would be responsible. It's very much possible to commission a ninja from a human artist and get something very pornographic on a famous celebrity and you would be held responsible for choosing to share it, you had the choice not to.
Megalomania: obsession with the exercise of power, especially in the domination of others
(What is silly is that DALLE2 won't let you edit a picture with a face in it, so you can't outpaint one of its own generations. But actually you can, if you crop it carefully.)
Software engineers who took the time to think about the ways that their work could be used.
How many times have we seen on HN people pleading with developers to think about the ethical dimensions of their work?
Just replace in scripts/txt2img.py:
- x_checked_image, has_nsfw_concept = check_safety(x_samples_ddim)
+ x_checked_image = x_samples_ddim
And be done with it.
- Communicate ahead of time; don't surprise maintainers with sweeping architectural changes or huge features no one wants or would like to review;
- Try to break up changes into logical units that can be understood and reviewed independently;
- Write useful and detailed commit messages (some bad examples: "Update attention.py", "various clean-ups, code now beautified");
- Don't sneak in anything unrelated to the PR; don't sneak anything unrelated into a commit;
- Absolutely don't use a code formatter to format the entire code base if the repo wasn't already using one. You can suggest that separately. And changes like that are best done by a trusted member.
Personally I blame juniors that insist on using the git CLI instead of a million visual GUIs to git. They get told just do a git commit, have no idea what they changed so can't produce a meaningful commit message, and because it's the command line good luck getting a visual representation of what you're committing.
I don't get the point here. I've never met a developer who knew about `git commit` but didn't know about `git diff` (sometimes I pointed out `git diff --cached`).
I don't think you can blame the CLI on this. It's about the lack of proofreading, in my opinion.
Maybe you work with a higher/better caliber of juniors than I do. At this point, I'm seriously contemplating being a mean dictator and forbidding commits outside of a dedicated GUI until they can prove their adeptness at using the CLI, which they should learn on their own time.
We Have Neon, who is making a PR.
We have Upstream, SD source.
We have basu, repo maintainer.
Neon pulled in upstream changes, and are trying to merge those into basu's repo AND then on top of that, apply neon edits to the code.
Ways to make this cleaner:
- Basu pulls in upstream changes, Neon just puts theirs on top. (Slow, you have to wait for Basu)
- Neon makes two PRs, one to pull in upstream changes, another for his edits on top of those changes. (More work, and coordination, but each PR is "one unit of work")
- Better commit messages: https://github.com/basujindal/stable-diffusion/pull/103#comm... <- notice the repeated and non-descript commit messages? That makes it hard for non-experts to cherry pick out the bits that are really relevant. (Fastest)
and this: https://www.aleksandrhovhannisyan.com/blog/atomic-git-commit...
That way your "large change" becomes a collection of small, easy to reason about and well explained changes.
It has an effect, but nothing like what's claimed in the submission title. On a 1070Ti (8GB), I managed to go up from 512576 to 576640.
Clone the original SD repo, which is what this code was built off of, and follow all the installation instructions:
https://github.com/CompVis/stable-diffusion
In that repo, replace the file ldm/modules/attention.py with this file:
https://raw.githubusercontent.com/neonsecret/stable-diffusio...
Now run a new prompt with a larger image. Note that the original model was trained on 512x512 and may lead to repetition especially if you try to increase both dimensions (this is mentioned in the SD readme) so just run with one dimension increased.
For example try the following example:
python scripts/txt2img.py --prompt "a person gardening, by claude monet" --ddim_steps 50 --seed 12000 --scale 9 --n_iter=1 --n_samples=1 --H=512 --W=1024 --skip_grid
I confirmed that if I run that command with the original attention.py, it fails due to lack of memory. With the new attention.py, it succeeds.
That said, this still uses 13GB of ram on my system.
I suppose you can check out the full repo with the updated code, which seems to have other changes, if you want to give that a try.
https://github.com/neonsecret/stable-diffusion/
I have already been using the original SD repo so I found benefit by just changing attention.py
https://github.com/basujindal/stable-diffusion/commit/47f878...
It's well engineered, maintainable and with decent installation process.
The branch in this PR[2] adds M1 Mac support with a one line patch and it runs faster than the CompVis version (1.5 iterations/sec vs 1.4 for CompVis on a 32 Gb M1 Max,
I highly recommend people switching to that version for the improved flexibility.
The MPS support issue for diffusers is here:
https://github.com/huggingface/diffusers/issues/292
…and it links to the relevant PyTorch issue here:
But the situation seems to be the same on the CompVis derived repos, right?
So this is no worse off, but with better engineered and faster code.
Having said that, the last comment [0] on the PyTorch issue gave me the idea of monkey patching the random functions. The supplied code assumes you’re always passing in a generator, which is not true in this case, but if you monkey patch the three rand/randn/randn_like functions to do nothing but swap out the device parameter for 'cpu' and then call to('ops') on the return value, it’s enough to get stable seed functionality for the CompVis derived repos without modifying their code, so I’m guessing it will probably work for diffusers as well.
Also, it’s probably a bug in the CompVis code, but even after you fix the random number generator, the very first run in a session uses an incorrect seed. The workaround is to generate an image once to throw away whenever you start a new session.
[0] https://github.com/pytorch/pytorch/issues/84288#issuecomment...
(Annoyingly I just went and made similar changes and was about to create a PR for them. But they have a fix for a "warm-up" issue I wasn't aware of too)
Copying this change fixed seeds on M1 for me.
Also, that’s only a partial fix that doesn't really work properly. It doesn’t affect img2img and it still gets things wrong on the first render. Since txt2img starts from scratch each time, that means you’re always getting an incorrect render, it just happens to be the same incorrect render each time.
[1] https://github.com/basujindal/stable-diffusion/pull/103/file...
> i) The model is being released under a Creative ML OpenRAIL-M license [https://huggingface.co/spaces/CompVis/stable-diffusion-licen...]. This is a permissive license that allows for commercial and non-commercial usage. This license is focused on ethical and legal use of the model as your responsibility and must accompany any distribution of the model. It must also be made available to end users of the model in any service on it.
[0] https://stability.ai/blog/stable-diffusion-public-release
https://m.youtube.com/watch?v=d_CgaHyA_n4&feature=emb_logo
And this is somehow relevant (possibly), as I kept the link open.
I use it directly from Python.
I can generate a 512x512 on my 10gb 3080 no problem (or three 384x384 at a time)
It measures memory usage as well.
[0] https://github.com/CompVis/stable-diffusion#stable-diffusion [1] https://github.com/basujindal/stable-diffusion#txt2img
If I was involved, this PR would be closed and conversation locked, simply because its a ton of formatting changes. The only "memory efficiency" changes are like 6 lines of python where the author uses `del`. Thats it. Everything else is formatting bullshit, and some other merge stuff. Yikes.
Also, some of those PR comments are written by kids or something. Wtf