ControlNET and Stable Diffusion: A Game Changer for AI Image Generation
bootcamp.uxdesign.cc
bootcamp.uxdesign.cc
As you can see, there is still quite a bit of flicker, I'm working to reduce that. But the consistency is much better compared to, say, img2img.
I'm hoping to ship a prototype this week.
I also had to specify what the outfit should be (I got a lot more discrepancies when I didn't do this from the outfit changing frame to frame). You can see that the outfit changes color in the second version, I bet you can get that to be even more consistent if you specify the color in the prompt too.
If you create a dreambooth model of a character, you can probably also get consistency of the face that way. In this case I didn't need to do this because I didn't care who I got, I just asked for an "average woman".
What could be of use here is a noise transformation layer that can use the same noise for every frame but transformed to match desired motion. For video conversion you could possibly extract motion vectors from successive frames to warp the noise.
I assume someone is working on this somewhere.
"Keeping the same seed wouldn't be helpful if you want the image to move." -- No, I'm using the same seed (and prompt). The image moves because ControlNet opens up another channel of input, in this case the pose data.
I spent hours trying to get a specific pose, hundreds of generations and seed changes, trying to dial in aspects of one vs another, taking aspects I like and clumsily patching together a Krita mash up then passing that back through img2img… only to get something kinda close.
One shot with canny or depth on Controlnet and I had exactly what I wanted.
Amazing how fast and easily it works. It’s been a WILD 6 months that this tech has been available.
I'm not sure how it all fits together but I'll try to figure it out. I have experience with Python development, git LFS, and I've been following research papers for years but this is my first attempt.
Is this all self-contained? Should I start elsewhere? I've been pretty intimidated by this stuff TBH, relying on hosted Midjourney / Dall-E for my artwork.
I notice the two repos have the same name. Do I merge the file-trees? How do they work together?
https://github.com/AUTOMATIC1111/stable-diffusion-webui
And then this https://www.reddit.com/r/StableDiffusion/comments/1167j0a/a_...
Step by step video tutorial https://youtu.be/vhqqmkTBMlU
Also fun fact, the human poses can be out-of-distribution and still works: https://twitter.com/toyxyz3/status/1626977005270102016
The use of eBsynth to stabilise leads to a pretty incredible result.
At least, we know it's capable of generating readable text, which SD sure isn't, and there are newer model papers out there like DiT that should beat latent diffusion on quality and speed.
Also, I'd hold my horses with calling latent diffusion "out of date". The Imagen paper notes that the results greatly scaled with text-encoder size and the model DeepFloyd (it's a team) are working on uses T5-XXL (the unet is also a decent bit larger, but the Imagen paper claims that this should have a much smaller influence). Latent diffusion might also be able to work with text if you scale up the text encoder.
I'm very skeptical of Google's "trust me, I have the hottest model, it just goes to another school". Muse might have nice fidelity and be fast, but the example images have worse visual quality than Imagen. I definitely wouldn't be surprised if transformers took over, but I don't think it's a foregone conclusion.
Your advanced model becomes a toy, like creating a Twitter clone on your local machine.
You then have to recreate all these things like control net, because I know for sure that this fine tuned control is way more important than a jump in base model quality.
This is why open source kicks ass. The SD people don't have to be the best, not with a community floating them.
Steamboat Willie remixes.