Diffusion with Offset Noise: Finetuning SD to generate very dark or light images
crosslabs.org
crosslabs.org
Just last week I saw ControlNet[1] which ads a lot more control.
Today I saw what corridor crew[2] did to stabilize the randomness when you want to make videos. Very exciting.
- User can install with a few clicks, no thinking about command line
- webUI has hover tooltips on everything, so user can most figure out what's going on without ever needing to touch documentation
- A1111 has a tab which can load a list of extensions from a github page
- Click to install, then refresh UI and extension just works
- Users are getting new incredibly powerful extensions every week or two - deforum lets you sequentially generate as many frames as you want and stitch them into a video, controlnet lets you copypaste features from a source image to your target image(s). Controlnet was added to A1111 ~2 weeks ago and is already integrated into the deforum tab so you can use both together.
Truly beautiful. I'd love to see more FOSS projects that felt so user friendly, generous with features, and rapid. Really fun to play with new cutting edge tech every couple weeks.
I was going to respond by saying it's not actually FOSS, but the author added an AGPL license last month: https://github.com/AUTOMATIC1111/stable-diffusion-webui/blob...
(I suspect they did not ask contributors to relicense, but hey)
That is, whatever color a region of the image has during denoising step 3, it will almost surely have that color at step 50, even if it makes no logical sense for the thing in that location to have that color.
This may not seem bad, but it's annoying when doing anything image-to-image, because regardless of the prompt you give it, the colors are "sticky".
If you have an image of an apple, and you use image-to-image with the prompt "an image of an orange", you will get a very reddish orange (in my experience at least).
I’m really new to all this and I’m learning new stuff about it every day!
I only understand a fraction of what this article says though :(
It's a way to cut just the components/styles/themes/patterns out from a model and apply them into other models.
So if I have a Disney characters checkpoint, but I really like this MakeGiantEyes checkpoint, if I can get it down to a MakeGiantEyes LoRA, I can apply that on top of my Disney Characters model which is already a custom trained set. It definitely does not always work, but when it does it's like magic. At a practical level, it's a model-modifier.
For example... Here is a Peter Griffin LoRA. https://civitai.com/models/13606/peter-griffin-lora or https://civitai.com/models/13763/thomas-the-train-i-lora
... It took me a minute to get those because I had to sort through a LOT things that would probably get me banned here. If anyone wanted to know nationalities were using SD more... It's Asians, hands down, all day long, and I think that's interesting.
EDIT: And if you were wondering what a Textual Inversion is vs a LoRA... Don't ask me! They're both model modifiers, but as I understand it, textual inversions are good for faces (which is why most of those are people, and they are kilobytes in size), and LoRAs aren't as good for faces specifically but better for themes.
TI exmaples (I couldn't use any of the million women... there are almost none that would be appropriate to post. Even though civitai does a good job of removing the nsfw posts of real people even with cloths on some are just still too much... Thirst is driving AI now) https://civitai.com/models/11039/ian-mckellen or https://civitai.com/models/8060/seu-madruga
https://civitai.com/models/8765/theovercomer8s-contrast-fix-...
I googled "wavelength of images" it doesn't seem like I am going in the right direction because it's about finding the wavelength of light from images rather than "wavelength of features" that this blog is talking about.
Pretty cool!
Will this cause us to hit walls at some point or actually exceed what a human can create?
We know that the biological brain does a lot of iterative refinement via recurrent processing (attractor dynamics), which is very similar to how diffusion works.
However, while prediction is also a core functionality of the brain, it's not really that you always auto-regressively generate word-by-word.