Generate images fast with SD 1.5 while typing on Gradio
twitter.com
twitter.com
Most people in the AI world don't understand that ML is like actual alchemy. You can merge models like they are chemicals. A friend of mine called it "a new chemistry of ideas" upon seeing many features in Automatic1111 (including model and token merges) used simultaneously to generate unique images.
Also, loras exist on a spectrum based on their dimensionality. Tiny loras should only be capable of relatively tiny changes. My guess is that this is a big lora, nearly the same size as the base checkpoint.
with or without finetuning? Also is there a practical motivation for creating them?
With, but it's still bonkers that it works so well
>Also is there a practical motivation for creating them?
You could get in-between model sizes (like 20b instead of 13b or 34b). Before better quantization it was useful for inference (if you are unlucky with vram size), but now I see this being useful only for training because you can't train on quants
Ehhhh…
LCMs are spooky black magic, I have no intuitions about them.
Now we find this way to skip to the end by building a model that learns the high dimensional curvature of the path that a diffusion process takes through space on its way to an acceptable image, and we just basically move the model along that path. That’s my naive understanding of LCM. Seems to good to be true, but it does work and it has a good theoretical basis too. Makes you wonder what is next? Will there be a single step network that can train on LCM to predict the final destination? LoL that would be pushing things too far..
You’d think that the fine tuning would make the LCM LoRA not work, but it does. Apparently the changes in weights introduced through even pretty heavy fine tuning does not wreck the transformations the LoRA needs to make in order to make LCM or other LoRA adaptations work.
To me this is alchemy.
So this is something that can be somewhat explained using not terribly handwavy mathematics. Picking hyperparameters on the other hand...
Pick eg. x -> sin(1/x) around zero and its derivatives.
The small modifications that you’re talking about are on the argument. These can lead to huuge changes in the values.
The stability is more likely due to the diffusive nature of the models and well executed trainings.
OTOH Gaussian kernels smoothen almost everything. Maybe it will be stable even with sin(1/x) as an “activation”.
Recalling the definition of exact differentiability is irrelevant.
Instead take the smallest interval that you can represent in fp32 not too far away from zero for example. Take few values in that infinite interval and check the behaviour of the said monstrous function.
This is a “trivial” example when studying eg. Distribution theory.
Said differently, you need to assess how smooth is the differential operator itself.
I’ve used it with my home GPU. Really fast which makes it more interactive and real-time.
I dream of the day we can sculpt entire 3D worlds with quick, tactile gestures. Or pure thought. This feels tangibly close.
Fantastic demo. Very much in the magic style of Johnny Chung Lee.
https://github.com/chengzeyi/stable-fast
Or AITemplate, and you are at 15FPS on a larger consumer GPU. 10 with a controlnet you can use for some motion consistency.
https://huggingface.co/collections/latent-consistency/latent...