Running Stable Diffusion on Your GPU with Less Than 10Gb of VRAM
constant.meiring.nz
constant.meiring.nz
It's simple and it works, using colab for processing but actually giving you a URL (ngrok-style) to open the pretty web ui in your browser.
I've been using that on-the-go when not at my PC and it's been working very well for me (after trying numerous other colab-dedicated repos, trying to fix them, and failing).
Additionally, you can have all your generated images sync to Google Drive automatically.
You can also buy Colab Pro and Colab Pro+, which have fewer limitations and faster GPUs.
I run it locally and can generate images with 50 steps in about 6 seconds per image, would it be faster for me to use Colab Free/Pro/Pro+?
1. Be central in the machine learning ecosystem. This has broad ripple effects, such as recruiting.
2. Doing things there means Google can track how you use machine learning. This can be used for everything from understanding trends in machine learning, to, again, robustly identifying individuals for recruiting efforts.
It seems like the cost is nominal at Google scale for what Google is getting. I suspect the pricing for the higher-end services is less a money-making scheme, as at some point, free is no longer sustainable (and if unlimited CPU were free, that would be prone to abuse / misuse / overuse / wasteful use). The amount of money Google makes there is nominal at Google scale.
Seeing that training etc. is much more memory intensive, and wanting to get faster results, I bought an RTX 3090, which has 24G of VRAM. However it maxes out at about 1024x512, only twice as many pixels. Observing the card with GPUZ, it never actually allocates more than 13.9G.
Using the lstein branch, I can't get above 896x512. Similarly, GPUZ shows allocated VRAM never reaches 14G. The interface isn't as good as the webui on hlky either - never mind the web interface, a bigger problem is it doesn't save all the parameters alongside generated images.
This is all running using Miniconda on Windows. On Linux it may be a different story, but my gaming PC is not dual-boot (yet).
My 1080 Ti gets around 2.5it/s with the k_lms sampler.
When I do batches, it slows down but not linearly; if 1x does 10 it/sec, 2x does about 6 it/sec. Batching is the other upside of more VRAM.
For my uses, the real benefit of having more VRAM is that you can generate more images simultaneously. My 3080 can generate only one 512x512 in 7 seconds but three 384x384 in that same timeframe. It’s allowed me to generate grids of hundreds of images in just a few minutes.
I bumped my system commit cap (increased paging file size) by 24G and now I can use all my VRAM.
Hope that helps somewhat. A 3090 should be more than enough for what you're doing, I'm stuck with P100s at best for me! (Cost :'( )
Best of luck! :D :)
Why the recommendation to stop using miniconda?
My only worry is that because the K80 is two GPUs on one board, that it might only utilize one of them, with only 12 GB of VRAM instead of both chips and all 24 GB.
5992 CUDA cores and 24 GB VRAM would be a pretty decent SD accelerator for only $150.
Craft Computing on Youtube has the best information from what I have seen so far. I don't like watching Youtube videos for information like this, but I understand why creators have moved to this medium in general. Linux should be much easier to configure for using the K80 to capacity.
> one minute per 512x512 @ 50 steps
> 1m20s to run 50 ddims on 512x512 vs 2080 ti in 12 seconds
You'll have to run the optimized model as well, since you can't connect the 2x 12GB together.
Any idea why this is?
Oh well, I've spent $90 on dinners and had far less fun than I will have with this video card when it arrives, so I can't say it was wasted money... and I can always just buy an M40 off eBay.
If RTX 3090 prices keep dropping through the floor, I may just bite the bullet and pick one up. I saw a ZOTAC on sale for $999 recently, which is $500 less than the launch MSRP of $1499 (which honestly is where it should have launched anyway... so far as I'm concerned, these cards only just now hit reasonable pricing).
>I think of it like a Python VM that just works where ever I place it.
That's called a virtualenv, which is a feature built into python. Miniconda is a thin wrapper around virtualenv (actual python packages) and the conda package format. If you're using an IDE it probably has virtualenv support baked in.
Personally I prefer to use python-poetry for managing virtual envs, but honestly just using the virtualenv command directly is not hard if you're already using conda from the CLI.
How it holds the cigarette with its little paw. Ehem, i mean, it's technically interesting, how the model correctly extrapolated, how this would look like..
- includes a nice GUI - txt2img and img2img - upscaling, face correction - many more
Bad enough I see trashy right-wing extremist shit over on the Stable Diffusion discord server zip past now and then.
I'll pass on that guide. Hopefully they grow the fuck up at some point.
I don't use anaconda so I created a new venv with python 3.10, installed the requirements as proposed, registered with hugging face and create the api key and run the provided source code.
Any way to improve the quality of the faces? Also how could I tune the parameters a bit ? (I'm not familiar with this AI stuff at all, I'm just a humble python programmer)
CUDA out of memory. Tried to allocate 3.00 GiB (GPU 0; 8.00 GiB total capacity; 5.62 GiB already allocated; 0 bytes free; 5.74 GiB reserved in total by PyTorch)
any ideas? I have a 3060ti with 8gb vram...with 448x448 I get:
CUDA out of memory. Tried to allocate 902.00 MiB (GPU 0; 8.00 GiB total capacity; 6.73 GiB already allocated; 0 bytes free; 6.86 GiB reserved in total by PyTorch)https://github.com/basujindal/stable-diffusion
https://github.com/neonsecret/stable-diffusion
Or the hlky webui, that is optimized too.
https://github.com/CompVis/stable-diffusion/issues/86#issuec...
What I did:
- scripts/txt2img.py, function - load_model_from_config, line - 63, change from: model.cuda() to model.cuda().half()
- removed invisible watermarking
- reduced n_samples to 1
- reduced resolution to 256x256
- removed sfw filter
Just can't get it to work and it's not producing an error message or anything that I could debug it with.
This comment: https://news.ycombinator.com/item?id=32710550 talks about running SD with 8GiB of VRAM and mentions needing to reduce this parameter to 1 to get it to output right.
My laptop takes about 6 seconds per iteration so it's significantly slower, but if you're willing to wait I bet you'll have a much easier time plugging more RAM into your system than adding VRAM.
For your case with 8 gb you shouldn’t need to do either of those things (run it all on gpu), just make sure you have batch size 1 and are using the fp16 version.
That's the world of running machine learning models for you. Why would anything ever work the first time right? Or at least the 10th time...
But one of the tricky parts with stable diffusion is that people are trying to get it to run on lighter hardware, which is basically another engineering problem where simple apis typically won't expose the kind of internals people want to mess around with.
Make sure you kill all python processes before restarting or some of your VRAM will be in use.
You can check with nvidia-smi how much ram is currently in use by what processes.
The not-optimized release works with my 2070 with 8 gb ram.
Also, you could try Visions of Chaos and use the Mode > Machine Learning > Text-to-Image > Stable Diffusion. It also has tons of other AI tools e.g. image-to-text captioning, diffusion model training, mandelbrot, music, and a ton more. The dev(s) push out updates almost every day.
Warning: You will first need to go through the 12 steps of Machine Learning setup first[0], then it will download 3-400GB of models since it has scripts for pretty much every latent diffusion out there, some of which e.g. Disco Diffusion I find to still give more interesting results and you can get much higher res on a 3060 Ti, plus you have a TON more parameters to play with, not to mention you can train your own models and load those in (which I've been doing the past few weeks using my photography to get away from using unlicensed imagery :)
[0] https://softology.pro/tutorials/tensorflow/tensorflow.htm
Particularly since cloud services are likely to be competitive and work for anyone.
M1 takes ~4.2s per iteration, 3.5 minutes per image [0].
RTX 3090 takes ~4.7s per image (all 50 iterations) [1].
[0] - https://wandb.ai/morgan/stable-diffusion/reports/Running-Sta...
[1] - trust me bro
I guess whomever ported that to M1 just haven’t implemented that method.
By default it's on, some forks have it turned off.
https://rom1504.github.io/clip-retrieval/?back=https%3A%2F%2...
So go here, turn off the safety filter and you can search to see what SD was trained on. I suspect that if you actually want the bush tit bird and donkeys, you'll want to use that instead.
Result: https://imgur.com/a/c1GM28U (NSFW)
Censored version would just replace anything NSFW with a picture of Rick Astley (for real [1]).
[0] - https://github.com/hlky/stable-diffusion
[1] - https://github.com/CompVis/stable-diffusion/issues/120
I use this on my Ubuntu 18 machine, works nicely on a GPU with 8GB VRAM.
As usual, some python dependency nonsense to sort out even with Anaconda, but pretty quick and easy to get up and running.
https://notes.datagenerator.eu/#Stable%20Diffusion%20install...
Bandwidth of dual channel DDR4-3600: 48 GB/s
Bandwidth of PCIe 4 x16: 26 GB/s
Bandiwdth of 3090 GDDR6X memory: 935.8 GB/s
Since neural network evaluation is usually bandwidth limited, it's possible that pushing the data through PCI-E from CPU to GPU is actually slower than doing the evaluation on CPU only for typical neural networks.https://www.microway.com/knowledge-center-articles/performan...
https://en.wikipedia.org/wiki/List_of_Nvidia_graphics_proces...
It worked perfectly fine, with the sole exception that the HDD LED was on solid the whole time, a single window took just over a literal half an hour to open, and loading a webpage took about 1-2 minutes.
But it worked.
https://github.com/AshleyYakeley/stable-diffusion-rocm
I had to make the one line change suggested in issue #3 to get it to run under 8GB.
radeontop suggests 4GB might work.
I also had to add this environment variable to make it work on my unsupported radeon 6600xt:
HSA_OVERRIDE_GFX_VERSION=10.3.0
It takes under two minutes per batch of 5 images with the --turbo option.
(Base OS is manjaro; using the distro's version of docker; not the flatpack docker package.)
If you don't have a GPU, paperspace will rent you an appropriate VM.
> Carbon Emitted (Power consumption x Time x Carbon produced based on location of power grid): 11250 kg CO2 eq.
That's ... Sobering.
Nobody seems especially bothered.
I now used parameters to drop the resolution to 256x256, and now it's running, but it's somehow broken. Every output image it produces is literally a green square.
On Linux with AMD `radeontop` is quite informative.
Does the image generator return some kind of seed that allows you to reproduce the result?
I am comically bad at getting it to generate what I want. e.g. "A furry watermelon" or "A dog flexing its biceps" just generates normal watermelons and normal dogs most of the time.
Any tips?
Edit: never mind this is the missing guide I had been looking for
It's also updated almost daily and tracks the latest features where possible.
I was also able to use the basic scripts to generate a few samples, pick one I liked, then used inpaint to expand the photo, masking out the original input so it wouldn't be altered.
After installing CUDA 11.7 and reinstalling torch I'm still facing:
> AssertionError: Torch not compiled with CUDA enabled
It could speed up calculations and significantly reduce memory requirements. I'd expect slightly worse results, though.
edit: also https://github.com/basujindal/stable-diffusion/pull/103
Seems like all of these projects are broken until you speak shibboleth by guessing at random python incantations. By this point, it's starting to feel intentional, like a way to mark you as part of an in-crowd, not a "L-User".
Unfortunately, I don't remember what I did. I did eventually get SD to work (though not in Docker, just as a normal python project). If I had been sober at the time, I probably would have given up. I know you need no greater than Python 3.9.
But, that's the point of the Miniconda dependency. By using Conda, the project can be setup locally with the Python version it expects without clobbering your local, system-level install.
It was still a bit of a pain to learn how to use Conda, but it worked out a little better than figuring it out on my own.
[1] https://docs.conda.io/en/latest/miniconda.html#linux-install...