High-performance image generation using Stable Diffusion in KerasCV
keras.io
keras.io
Edit: there seems to be a more "full" version of the same work available here, made by one of the authors of the submission article: https://github.com/divamgupta/stable-diffusion-tensorflow
This is by far the most popular and active right now: https://github.com/AUTOMATIC1111/stable-diffusion-webui
There's also this fork of AUTOMATIC1111's fork, which also has a Colab notebook ready to run, and it's way, way faster than the KerasCV version: https://github.com/TheLastBen/fast-stable-diffusion
(It also has many, many more options and some nice, user-friendly GUIs. It's the best version for Google Colab!)
They sure do. InvokeAI is a fork of the original repo CompVis/stable-diffusion and thus shares its fork counter. Those 4.1k forks are coming from CompVis/stable-diffusion, not InvokeAI.
Meanwhile AUTOMATIC1111/stable-diffusion-webui is not a fork itself, and has 511 forks.
Thanks for the correction.
Any idea on how to count forks of a downstream fork? If anyone would know... :)
The only command line argument I'm using is --lowvram, and usually generate pictures at the default settings at 512x512 image size.
You can see all the command line arguments and what they do here: https://github.com/AUTOMATIC1111/stable-diffusion-webui/wiki...
While technically the most popular, I wouldn't call it "by far". This one is a very close second (500 vs 580 forks): https://github.com/sd-webui/stable-diffusion-webui/tree/dev
AUTOMATIC has the attention of most devs nowadays. When you see any new ideas come up, they usually appear in AUTOMATIC's fork first.
Short version: Attention works as a matrix multiply that looks like this: s(QK)V where QK is a large matrix but Q,K,V and the result are all small. You can break Q and V into horizontal strips. Then the result is the vertical concatenation of:
s(Q1*K)*V1
s(Q2*K)*V2
s(Q3*K)*V3
...
s(QN*K)*VN
Since you're reusing the memory for the computation of each block you can get away with much less simultaneous RAM use.You can do something like that but it's far from optimal.
From memory consumption perspective, the right way to do it, is to never materialize the intermediate matrices.
You can do it, by using a customop, that compute att = scaledAttention(Q,K,V) and the gradient dQ,dK,dV = scaledAttentionBackward(Q,K,V,att,datt)
The memory needed for these ops is the memory to store Q,K,V,attn,dQ,dK,dV,dattn + extra temporary memory.
When you do the work to minimize memory consumption, this extra temporary memory is really small : 6attention_horizon^2number_of_core_running_in_parallel numbers.
But even though there is not much re computation, this kernel won't run as fast due to the pattern of memory access, unless you spend some time manually optimizing it.
The place to do it is at the level of the autodiff framework aka tensorflow or pytorch, with low level c++/cuda code.
Anybody can write some custom kernel, but deploying, maintaining them and distributing them is a nightmare. So the only people that could and should have done it, are the tensorflow or pytorch guys.
In fact they probably have, but it's considered a strategic advantage and reserved for internal use only.
The mere mortals like us, have to use some workarounds (splitting matrices, Kheops, gradient checkpointing... ) to not be too much penalized by the limited ops of the out of the box autodiff frameworks like tensorflow or torch.
Either way, getting 3 images for 25 iterations under 10 seconds (quick Colab test, which is where I've taken to comparing these things) is just ridiculously faster.
PyTorch is now quite a bit more popular than Keras in research-type code (except when it comes from Google) so I don't know if these enhancements will get ported. This port was done by people working on Keras which is kind of telling - there isn't a lot of outside interest.
We got a Keras version made by Divam Gupta very quickly after Stable Diffusion was released.
Is he not an outsider?
https://github.com/divamgupta/stable-diffusion-tensorflow
Now they are working together. That may be “telling” to you but I’m not sure why that should cast a negative light on Keras, really.
It’s hard to understand how that can say anything negative about the popularity of Keras among outsiders.
The OP was wondering whether additional enhancements will also be ported and that's what I responding to. It's simply much less likely that a new paper will get a Keras implementation than a PyTorch implementation.
Because when Keras creators learned about the port done by what you define as an outsider they thought that it was cool - and that it would be nice to make it part of the KerasCV toolbox.
What answer did you expect?
You understand that it was "an outsider" who ported Stable Diffusion to Keras, right?
Timeline:
Aug 22, 2022 [day 1 of the SD era] Public release of stable diffusion
Sep 17, 2022 [day 27 of the SD era] Keras outsider Divam Gupta announces a Keras port
https://twitter.com/divamgupta/status/1571234504320208897
"Stable Diffusion implemented using @Tensorflow and #Keras. [...] Thanks @fchollet and team for building this amazing framework which makes it easy to implement a model like Stable Diffusion."
Sep 19, 2022 [day 29 of the SD era, 2 days after port announcement] François Chollet publishes a Twitter thread about the port (and his own improvements on a fork - two dozen commits dated Sep 18-21 will be pushed upstream)
https://twitter.com/fchollet/status/1571874757582389250
"Huge thanks to @divamgupta for creating this port! This is top-quality work that will benefit everyone doing creative AI. I'm always amazed by the velocity of the open-source community"
Sep 25, 2022 [day 35 of the SD era, 8 days after port] Divam Gupta talks about the upcoming KerasCV integration
https://twitter.com/divamgupta/status/1574050880000770050
"Last week I implemented Stable Diffusion using Keras / Tensorflow. Now its almost integrated in KerasCV thanks to @fchollet and team. This is the power of open source collaboration."
Sep 27, 2022 [day 37 of the SD era, 10 days after port announcement] Official release of the Stable Diffusion implementation in KerasCV - and TFA (dated Sep 25)
https://twitter.com/fchollet/status/1574782176633430017
"Stable Diffusion is now available directly in KerasCV! [...] Many thanks to all those who made this implementation possible, in particular @divamgupta @luke_wood_ml and of course the creators of the original Stable Diffusion models!"
The benchmark in the article uses mixed precision (and equivalent generation settings) for both implementations, it's a fair benchmark.
In the latest StackOverflow global developer survey, TensorFlow had 50% more users than PyTorch.
It also doesn't help that PyTorch has its own discussion forum [1] where most pytorch questions end up.
It is a bit sad if this is just a closed software issue that cannot be fixed :(
The prompt was "A cute otter in a rainbow whirlpool holding shells, watercolor"
Seems like the otter should be holding shells, the way a normal human parses it.
The tool showed the otter 'in holding-shells', which are shells that hold otters apparently. Also some random shells strewn about, as the technique is sensitive to spurious detail sprouting up from single words.
Until the tool permits some kind of syntactic diagramming or so forth, we'll not be able to control for this.
Just the other day here, I saw a picture of a fork and some plastic mushrooms. The prompt was 'plastic eating mushrooms' which was ambiguous even to humans. The tool chose to illustrate the subclass of mushrooms 'eating-mushrooms' (as opposed to poison mushrooms or decorative mushrooms I suppose) made of plastic.
When we're playing around this can seem whimsical and artistic. But a graphic designer might want some semblance of control over the process.
Not sure how a solution would work.
"
Compositional generation. Our method can compose multiple diffusion models during inference and generate images containing all the concepts described in the inputs without further training"
The model is loaded from Huggingface during the instantiation of the stable diffusion class. It is loaded as an H5 file which I believe is unique to Keras[0]. I don't have any experience with Keras so I can't say if that is good or bad. I wanted to see where they were getting the weights as the blog post didn't demonstrate an explicit loading function/call like Pytorch.
Gonna run it and see... although I have like 40GB of stable diffusion weights on my computer now.
[0] https://github.com/keras-team/keras-cv/blob/master/keras_cv/...
I guess I should give it a try.
I tried some and found no major differences after 16 steps or so with given random seed.
[1] https://github.com/CompVis/stable-diffusion/
[2] https://github.com/divamgupta/stable-diffusion-tensorflow/
Something like global optimization has been done in pytorch, here's a blog about it: https://www.photoroom.com/tech/stable-diffusion-25-percent-f...
Mixed precision seems pretty much default looking at a few Stable Diffusion notebooks.
More intriguing, there's also a more local optimization that makes pytorch faster: https://www.photoroom.com/tech/stable-diffusion-100-percent-...
Unless it's already there, that last one would be interesting to add to keras.
All in all this machine learning ecosystem is wild, as a software dev, things like cache locality and preferring computation over memory access are basic optimizations, yet in machine learning it seems wildly disregarded, I've seen models happily swapping between gpu and system memory to do numpy calculations.
Hopefully stable diffusion changes things, the work towards optimizations is there, it just seems often disregarded. As stable diffusion is one popular open model that, when optimized, can be run locally (and not as saas, where you just add extra compute power, which seems cheaper than engineers) and has a lot of enthusiasm behind it, it might just be the spark that makes optimization sexy again.
A problem I see is that a lot of times everything works fine on rocm+hip, but since nvidia dominates the machine learning market (and thus most researches run nvidia), most forks don't bother checking and just advertise compatibility with nvidia and sometimes apple M1.
Problem is, AMD GPUs are much cheaper!
AMD seems to be popular among gamers on budget and the budget cards often don't have the VRAM required by default. So, AMD seems to be in this weird place where the people who can make it work don't care.
Are they? I believe Nvidia (consumer) gpus have better price/performance than amd for AI.
Here I'm considering that the main constraint is VRAM and while stable diffusion now runs even on GPUs with 2GB RAM, there's always new developments that require more VRAM (for example, Dreambooth requires 12GB as of today)
Which repo or build are you using BTW, is it the one related to this readme?
https://github.com/magnusviri/stable-diffusion/blob/main/REA...
Yes, this one. However it was like a month ago I think, so speeds might have improved. I'm getting ~2.2s/step with another implementation: https://news.ycombinator.com/item?id=33006447
I am also wondering, do you follow the general advice of 1 iteration and 1 sample, for example:
--n_samples 1 --n_iter 1 (when referencing commands using txt2img.py)
I figure you could wait a bit for things to process going further, but curious just if you're getting results like that with higher sample/iter settings.
This is an example of the original file: https://github.com/magnusviri/stable-diffusion/blob/79ac0f34...
Which seems to have been renamed, and cleaned up a bit here: https://github.com/magnusviri/stable-diffusion/blob/main/doc...
However, per the note on the magnusviri repo, the following repo should be used for a stable set of this SD Toolkit: https://github.com/invoke-ai/InvokeAI
with instructions here https://github.com/invoke-ai/InvokeAI/blob/main/docs/install...
https://reddit.com/r/StableDiffusion/comments/xbo3y7/oneclic...
Maybe they use lower parameters.
edit:
50 steps at 256x256 resolution took 55 seconds.
50 steps at 768x768 resolution took 8 min, exactly.
PS: my Macbook Air is modified with thermal pads, it takes a bit longer to start throttling than usual. Either way, it's very dependent on the ambient temperature.
It seems like using high precision is useful for training, but if not training, why not just use float16 weights and save the memory?
If you really just want to save memory, there's plenty of other low hanging fruit. It's just not a priority for most devs since mid tier GPUs start at 10GB whereas a typical model only has 0.5GB weights. Activations and intermediate calculations use way more memory.
Has anyone got a suggestion on how to fine tune this model?