How to Build Your Own AI-Generated Images with ControlNet and Stable Diffusion
datature.io
datature.io
Was getting good results with just SD and occassional masking but it would take hours and hours to hone in and composite a complex scene with specific requirements & shapes (with most of the work spent currating the best outputs and then blending them into a scene with Gimp/Inkscape).
Masking is unintuitive compared to the scribble which gets similar effect; no need to paint masks (which is disruptive to the natural process of 'drawing' IMO) instead just make a general black and white outline of your scene. Simply dial up/down the conditioning strength to have it more tightly or fuzzily follow that outline.
You can also use Gimp's Threshold or Inkscape Trace Bitmap tool to get a decent black & white outline from an existing bitmap to expedite the scribble procedure.
Reminds me of Fireworks which Adobe killed off (after putting out a decent update or two to be fair) which used PNGs for layers and meta ala PSD format.
But its more analogous to a 3D modeller suite like Blender or Maya but with theoretical feature such that you could take a rendered output image and dragndrop it back into the 3D viewport and have it restore all the exact render settings you used instantly back. That would be handy!
This is super powerful for example visualizing the renovation of an apartment room or house exterior.
In this post, we just seek to showcase the fastest way to do it - and how augmentation may potentially help vary the position!
Note that all of the images in those comfy tutorials (except for images of the UI itself) can be dragndropped into ComfyUI and you'll get the entire node layout you can use to understand how it works.
Another good resource is civit.ai and specifically look for images that have a comfy UI embedded metadata. I made a feature request that they create a tag for uploaders to flag comfyUI pngs but not sure if they've added that yet. Or caroose Reddit or Discord for people sharing PNGs with comfy embeds.
Trying out different models (also avail from civit) is a good way to get an understanding of how swapping out models affects performance and the results. I've been abusing Absolutereality (v1.81) + More Details LORA because its just so damn fast and the results are great for almost any requirement I throw at it. AI moves so fast but I don't even bother updating the models anymore there is just so much potential in the models we already have; more pay off would be mastering other techniques like the depth map Control Net.
I would say that above all extensive familiarity with an image editor like Photoshop, Gimp, or Krita - will get you the most mileage particularly if you have specific needs beyond just fun and concepting. AI art makes artists better, people who struggle with image editing will struggle to maximize this new tech just as people who struggle with code will have issues maintaining the code Copilot or ChatGPT is spitting out (versus a coder who will refactor and fine tune before integrating to the rest of their application).
It's still a bit rough around the corners, and I haven't properly launched it yet, but if you want to play with ControlNets, pre-processors, IP adapters, and all those various SD technologies, it's a pretty fun tool ! I personally use for real-time scribble to image, things like this :)
(will post that properly on HN in a few days / week I think, when early feedback will have been properly addressed)
I barely got it working in that early alpha but it was super helpful for me as a reference. I'll give it another go now that it's further along, it seemed very promising and I liked your workflow approach
Prompts seem to be a new type of camera, lens or paintbrush.
Style generality is frequently lost in fine-tuned models. The original dreambooth tried to get around this by generating lots of images of the class to retain generality, but it's time intensive to generate all the extra images (and ideally do some QC on them) and train on them too, so it's not often done.
Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language.
I guess my point is to ask whether SD is worth bothering with at this time when DALL E and Imagen and possibly others are just on the brink of becoming mainstream and just going to get better and better. Clunking together something with SD seems unnecessary when you can generate more results, better results, in a faster way, with less requirements, and without the steep learning curve, by using other methods.
Depends if you value this kind of freedom in life.
You can do some things such as colorizing black and white images with the Recolor model.
Stable Diffusion is an open model that you can run locally on your own computer without anyone’s permission. Dall-E is a closed model that runs on OpenAI’s very expensive server farm, and they can change how it works and what it costs whenever they please.
Right now AI is in the Uber-style expansion phase where the service is practically given away to conquer market share. Once the hypergrowth is over, OpenAI will start raising their prices just like Uber did.
Plus a million other tools that the community has made for it, like ControlNet or things like AnimateDiff to create videos. I can also easily create all kinds of scripts and workflows.
Are you getting around that somehow? Even if it'll let me generate 36 images per half hour (which seems like it's probably lower than that) I can only generate 6k in 2 weeks prompting 24/7. I'm not scrutinizing your numbers I'm more hoping I'm missing some way to not have to be capped. I already pay for GPT+
No, you know when a beginner generated an image in Stable Diffusion. With enough skill and attention, you will not.
Sure, there is a learning curve and it takes more time to get to a good result. But in turn, it gives you control far beyond what the competition can offer.
Give it a go with invokeAI - you can create images that I guarantee you wouldn’t know were generated. Like anything (photography included) it’s a skill.
Examples:
- https://civitai.com/images/2862100
- https://civitai.com/images/2339666
- https://civitai.com/images/2846876More than that though: I use SDXL quite a bit for fun, and while I like it and it can be very good, it's still prone to getting stuck in a David Cronenberg mode for reasons I can't solve.
I cant change anything on DALL E, I can just take the input or change the prompt.
Also it is a centralized service that can be shut down, modified, censored or become very expensive at any time.
If you see a part of the scene that looks weird (and you know what it should be) add it to your prompt. For example, if you want "photo of a jungle in South America", and the foliage looks weird, add something like "with lush trees and ferns".
I also recommend a good photorealistic base model, like RealVis XL.
In my experience its like DALL E but straight up better, more customizable, and local. And thats before you start trying finetunes and LORAs.
Other UIs will do SDXL, but every one I tried is terrible without all those default fooocus augmentations.
It has plenty of other advantages, but you can't tell it "make me a cute illustration of a 2 year old girl with Blaze from Blaze and the Monster Machines on a birthday cake with a large 2 candle on it."
DALL E will nail that, more or less. SDXL very much won't.
- https://ibb.co/k0NCWG7 - https://ibb.co/Vm3GZcR - https://ibb.co/bvSC4w3 - https://ibb.co/VqSdYbZ
I'm surprised that it didn't complain about copyrighted characters, it tends to do that a lot for me.
More cherrypicking and messing with styles is getting closer, but nothing like Dall-E's first try I'm sure.
You don't. People think they do, but they don't.
And, indeed, someone has:
Dalle 3 is super good, but lacks the creative control controlnets and ip-adapter provide. So for instance afaik there is no way to perform style transfers, or ’paint a van gogh portrait over my pencil sketch’.
Both are good currently but at different things.
”Prompt engineering” is and will be total bs. Dalle3/chatgpt provides the actual workflow we want where we describe to the intelligent agent (chatpgt) what we want and it worries over the accidental-complexity-intricasies of the clip model itself.
SD is worth bothering with because it's open, you can run and extend it yourself.
>Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language.
Hyper-realistic, but is it what you want from it? Are you able to guide it into doing exactly what you want? If you have such requirements that just a natural language prompt is enough and is somehow faster than sketching and providing references, of course use it. I'm not so lucky, I don't get what I want from it, and no amount of prompt understanding will make it easier. Although SD/SDXL doesn't pass the quality bar either, not because it's not "detailed" or "hyper-realistic" enough, but because it doesn't pay attention to the things that should be prioritized, like linework or lighting. Neither does any other model. Controlnets and LoRAs alone aren't sufficient for controllability either, mostly because it's too small to understand high-level concepts. So I don't use anything.
It's pricey to get a windows machine + GPU and the cloud options seem a bit more limited and add up quickly too, but it is amazing tech.
Here is a colab link to open comfyUI
https://github.com/FurkanGozukara/Stable-Diffusion/blob/main...
It’s the best argument against “AI generated images are just collages”.
LLMs + Diffusors are super charged when using techniques like constraints, controlnet, regional prompting, and related techniques.
If you're serious about learning how to use these tools, it's far more affordable to rent a GPU in the cloud. Google Colab even has some free tiers with limited access to GPUs that are significantly more powerful than what you would normally put in a desktop.
I just want to create my own images like Dall-e 2 but using my face instead.
> P.S. As pointed out by a fellow HackerNews reader, we clearly forgot to include our code snippet for ControlNet in the article.
No other code snippet besides the one added in response uses Canny, at least so far as I can see.