Like diffusion but faster: The Paella model for fast image generation
deeplearning.ai
deeplearning.ai
My 3060 can generate a 256x256 8 step image in 0.5 seconds, no A100 needed. A 3090 is double the performance of a 3060 at 512x512, and an A100 is 50% faster than a 3090...
If you have access to an high end consumer GPU (4090) you can generate 512x512 images in less than a second, it's reached the point that you can increase the batch size and have it show 2-4 images per prompt without adversely affecting your workflow.
Too bad SD1.5 is too small* and we'll require models with more parameters if we want a true general purpose image model. If SD1.5 was the end-game, we'd have truly instant high res image generation in just a couple more generations of GPUs, think generating images in real time as you type the prompt, or have sliders that affect the strength of certain tokens and see the effects in real time, etc. Tho I heard that SDXL is actually faster for higher resolutions (>1024x1024) due to removing attention on the first layer, making it scale better with resolution even tho SDXL has 4x the parameter size.
* Current SD1.5 models that can generate consistent high quality images have been fine-tuned and merged so many times that a lot of general knowledge has been lost, e.g. they can be great at generating landscapes, but lacking in generating humans, or they can be very good at a certain style like comics but can do comic style only and lose the ability to generate more dynamic face variations, etc.
It's fairly impressive to me what the community has made possible with SD1.5. Sure on a vanilla task something like Dall-E 2 generally performs better, but with some tweaking you can easily beat out Dall-E on a home gaming PC.
The fact that you can fine-tune SD1.5 on a 4090 is incredible to me.
Given how much powerful AI is locked behind fees and private APIs it's refreshing to see so much cool stuff coming out of the OSS world again. Best of all is it's not being driven exclusively by people with an ML background, but moreso curious amateurs. It really brings me back to a time when playing around with software/the web felt exciting.
I would be more interested to see Paella vs SD running on a ML compiler framework, like TVM or AITemplate. Maybe one or the other is more amenable to optimization.
The trick is to simply not do any calculations till the last possible moment - ie. the time the program tries to convert the finished image to a jpeg. Only at that point do you compile the graph and run the actual computation on the GPU.
You then also cache the graph, so that the compilation step can be avoided if the program tries to do the same computation again with different data.
That approach is limited though. AITemplate and TVM take a looong time to compile and produce standalone executable files, hence the gains are much larger than torch triton.
EDIT: If I understand correctly these libraries target deployment performance, while torch.compile is also/mostly for training performance?
- Torch 2.0 only supports static inputs. In actual usage scenarios, this means frequent lengthy recompiles.
- Eventually, these recompiles will overload the compilation cache and torch.compile will stop functioning.
- Some common augmentations (like TomeSD) break compilation, force recompiles, make compilation take forever, or kill the performance gains.
- There are othdr miscellaneous bugs, like compilation freezing the Python thread and causing networking timeouts in web UIs, or errors with embeddings.
- Dynamic input in Torch 2.1 nightly fixes many of these issues, but was only maybe working a week ago? See https://github.com/pytorch/pytorch/issues/101228#issuecommen...
- TVM and AITemplate have massive performance gains. ~2x or more for AIT, not sure about an exact number for TVM.
- AIT supported dynamic input before torch.compile did, and requires no recompilation after the initial compile. Also, weights (models and LORAs) can be swapped out without a recompile.
- TVM supports very performant Vulkan inference, which would massively expand hardware compatibility.
Note that the popular SD Web UIs don't support any of this, with two exceptions I know of: VoltaML (with WIP AIT support) and the Windows DirectML fork of A1111 (which uses optimized ONNX models, I think). There is about 0% chance of ML compilation support in A1111, and the HF diffusers UIs are less bleeding edge and performance/compatibility focused.
And yes, triton torch.compile is aimed at training. There is an alternative backend (Hidet) that explicitly targets inference, but it does not work with Stable Diffusion yet.
You say that as if it was a bad thing, but it's actually good: GPU memory being a limiting factor means that general knowledge is mostly overhead, it is much better to have 50 specialized model (that you can all store on disk for cheap) that each takes 5 time less GPU memory than a big general model that you'll constantly under-use but still have to load entierly in the GPU memory. And it's even more true for LLMs.
One of my dad's preferred anecdotes about how much computers sped up in his career, was the number of digits of pi that the company mainframe could compute.
He was born in '39.
And now I can generate images from descriptions faster than I can give those descriptions.
At this rate, websites will be replaced with image generators and LLMs, and the loading speed won't change.
But to continue their story, a lot of us used to play games at a 256x256 resolution.
So in the grand scheme of "things improve quickly, by a lot" it very much applies.
Given that Paella uses tokens instead of the source image, I wonder if the results will have a (human- or machine-) detectable "style" to them.
With automatic1111 you can get around this by upscaling then inpainting the spots you want more detail and specifying a specific prompt for that particular area.
The model can't just work on arbitrary image sizes because the model was trained with a fixed number of input and output neurons. For example, 512x512 is Stable Diffusion's "native size." However, there are tricks to work around this.
Diffusion models work by predicting image noise, which is then subtracted from the image iteratively until you get a result that matches the prompt. Stable Diffusion specifically has the following architectural features:
- A Variational Autoencoder (VAE) layer that encodes the 512x512 input into a 128x128 latent space[0]
- Three cross-attention blocks that take the encoded text prompt and input latent-space image, and output a downscaled image to the next layer
- A simpler downscaling block that just has a linear and convolutional layer
- Skip connections between the last four downscaling blocks and corresponding upscaling blocks that do the opposite, in the opposite order (e.g. simple upscale, then three cross-attention blocks).
- The aforementioned opposite blocks (upscale + cross-attn upscale)
- VAE decoder that goes from latent space back to a 512x512 output
At the end of this process you get what the combined model thinks is noise in the image according to the prompt you gave it. You then subtract the noise and repeat for a certain number of iterations until done. So obviously, if you wanted a smaller image, you could crop the input and output at each iteration so that the model can only draw in the 'center'.
Larger images are a bit trickier, you have to feed the image through in halves and then merge the noise predictions together before subtracting. This of course has limitations: since the model is looking at only half the image, there's nothing to steer the overall process, so it will draw things that look locally coherent but make no sense globally[1].
I suspect - as in, I'm totally guessing here - that we might be able to fix that by also running the diffusion process on a downscaled version of the image and then scaling the noise prediction back up to average with the other outputs. As far as I'm aware no SD frontends do this. But if that worked you could build up a resolution pyramid of models at different sizes taking fragments of the image and working together to denoise the image. If you were training from scratch you could even add scale and position information to the condition vector so the model can learn what image features should exist at what sizes.
[0] Think of this like if every pixel of the latent-space image was, instead of RGB, four different channels worth of information about the distribution of pixels in the color-space image. This compresses the image so that the U-Net part of the model can be architecturally simpler - in fact, lots of machine learning research is finding new ways to compress data into a smaller amount of input neurons.
[1] Moreso than diffusion models normally do
From the Paella paper[2]: "Our proposal builds on the two-stage paradigm introduced by Esser et al. and consists of a Vector-quantized Generative Adversarial Network (VQGAN) for projecting the high dimensional images into a lower-dimensional latent space... [w]e use a pretrained VQGAN with an f=4 compression and a base resolution of 256×256×3, mapping the image to a latent resolution of 64×64indices." After training, in describing their token predictor architecture: "Our architecture consists of a U-Net-style encoder-decoder structure based on residual blocks,employing convolutional[sic] and attention in both, the encoder and decoder pathways."
U-Net, of course, is a convolutional neural network architecture. [3]. The "down" and "up" encoder/decoder blocks in the Paella code are batch-normed CNN layers. [4]
[1] https://arxiv.org/pdf/2012.09841.pdf [2] https://arxiv.org/pdf/2211.07292.pdf [3] https://arxiv.org/abs/1505.04597 [4] https://github.com/dome272/Paella/blob/main/src/modules.py#L...
> Running on an Nvidia A100 GPU, Paella took 0.5 seconds to produce a 256x256-pixel image in eight steps, while Stable Diffusion took 3.2 seconds
Using the latest methods (torch 2.0 compile, improved schedulers) stable diffusion only takes about 1 second to generate a 512x512 image on an a100 gpu. A 256x256 image 1/4 the size presumably takes less than half that time.
So the corrected title is “Like diffusion but slightly slower and lower quality.”
[0]:https://www.reddit.com/r/StableDiffusion/comments/z3m97e/min...