Show HN: Scribble Diffusion – Turn your sketch into a refined image using AI
scribblediffusion.com
scribblediffusion.com
But for people who like it, the ControlNet scribble model (and the other ControlNet models, depth-map based, pose control, edge detection, etc.) [0] are supported in the ControlNet extension [1] to the A1111 Stable Diffusion Web UI [2], and probably similar extensions for other popular stable diffusion UIs. Should work in any current browser, and at least the A1111 UI, with ControlNet models, works on machines with as little as 4GB VRAM.
[0] home repo: https://huggingface.co/lllyasviel/ControlNet but for WebUI you probably want the ones linked from the readme of the WebUI ControlNet extensions [1]
[1] https://github.com/Mikubill/sd-webui-controlnet (EDIT: even if you aren’t using the A1111 WebUI, this repo has a nice set of examples of what each of the ControlNet models does, so it may be worth checking out.)
Edit: nvm, this particular demo does require you to type in positive prompt.
With more playing, I’d say this probably doesn’t have a canned positive prompt, just the user input.
Free and open source [3].
Bonus: it works on Firefox (unless you’re using private mode — because Firefox doesn’t make IndexedDb available to web apps in private mode and things end up breaking)
[1] https://tinybots.net/artbot/controlnet
[2] https://tinybots.net/artbot/draw
[3] https://github.com/daveschumaker/artbot-for-stable-diffusion
Thanks for your kind words and feedback. I love seeing all these links to your scribbles.
I'm an engineer at Replicate, which is a place to run ML models in the cloud. [0] We built Scribble Diffusion as an open-source app [1] to demonstrate how to use Replicate.
This is all built on ControlNet [2], a brilliant technique by Lvmin Zhang and Maneesh Agrawala [3] for conditioning diffusion models on additional sources, in this case human scribbles. It also allows for controlling Stable Diffusion using other inputs like pose estimation, edge detection, and depth maps.
ControlNet has only existed for three weeks, but people are already doing all kinds of cool stuff with it, like an app [4] that lets you pose a stick figure and generate DreamBooth images that match the pose. There are already a bunch of models [5] on Replicate that build on it.
I see a few bits of feedback here about issues with the Scribble Diffusion UI, and I'm tracking them on the GitHub repo. If you want to help out, please feel free to open an issue or pull request.
[1] https://github.com/replicate/scribble-diffusion
[2] https://github.com/lllyasviel/ControlNet
[3] https://arxiv.org/abs/2302.05543
[4] https://twitter.com/dannypostmaa/status/1630442372206133248
“a goofy owl”: https://scribblediffusion.com/scribbles/oymg4kadgvezxppvwkf5...
“a goofy owl, realistic photograph, depth of field, 4k, HDR”: https://scribblediffusion.com/scribbles/va5l24amjnb55g62renz...
“a goody owl, pointillism”: https://scribblediffusion.com/scribbles/5dfnru4f6zguphjvvdrl...
It seems this does not work on Firefox? I could only draw on about half the canvas and it was pretty buggy. Dont support the chrome monoculture!
https://github.com/vinothpandian/react-sketch-canvas/pull/11...
Sorry about the trouble. The Firefox incompatibility is the result of a bug in the underlying npm package we're using to render the drawing tool and canvas.
The issue is being tracked here: https://github.com/replicate/scribble-diffusion/issues/17#is...
We may need to wait for a fix to that, or consider swapping out the package we use for scribbling on a canvas.
"hummingbird drinking from tulip" disaster https://scribblediffusion.com/scribbles/n2ekqcs7vnegdcezcdvg...
"hummingbird on left drinking from tulip on right" perfect https://scribblediffusion.com/scribbles/ri5y2kzanzcs7dvxhlgy...
[1] “treasure map in the style of lord of the rings maps” https://scribblediffusion.com/scribbles/t36as45npjapxfekcuuy...
Uploaded photos Examples: https://imgur.com/a/jeWgRvH Website: https://crappydrawings.ai
Now trying to decide if this is healthy for kids.... who may lose motivation to ever pursue detailed art and drawing because "AI can do it for me based on rough sketch". But at same time, it may motivate them to make more rough drawings with interesting ideas.
I can see storyboarding for film will make use of this. Although having the characters and background remain consistent from scene to scene is something I don't know how can be done, if every submission is a roll of dice.
If you already have your own images, you can use the Replicate model directly: https://replicate.com/jagilley/controlnet-scribble -- you can upload your image using the Replicate web UI or do it programmatically using the model's HTTP API.
I don't quite understand why some image generators are free and others charge credits, assuming the non-free ones aren't just goldrushing.
Is it just free until oops it gets too popular?
Of course running on rented/hosted GPU's, it's a simpler, but much more expensive story — basically however much you're paying for GPU instances to run Stable Diffusion divided by how many images you generate. :)
[0]: https://scribblediffusion.com/scribbles/elun6gwkxrcr7eqy5jao... [1]: https://scribblediffusion.com/scribbles/tpsty6qcxjfrxbxz3n6b...
I use that prompt because it's something my preschooler draws on a daily basis with no qualms at all, but can often be hard for generative AI to imagine. So it's getting close to the imaginative abilities of a 3 year old! Halfway there.
https://i.imgur.com/08o9zkG.jpg
https://i.imgur.com/KkBtTyd.jpg
Eventually I did get something which looked like a partial success, but with low resolution and not something I’d consider appealing:
https://i.imgur.com/SJWkwBp.jpg
https://i.imgur.com/BpSUD4j.jpg
This is not a dig on the author. I have yet to see a simple prompt give a good result with Stable Diffusion. Are there examples of it?
are you generally an unlucky person? :D
Well, it guessed which character I drew...
https://scribblediffusion.com/scribbles/vaoxhqknfrdxpb2osym2...
Not bad. Even came close with the headband. But it figured her ear was an AirPod. (Maybe a small persocom earpod? Is this the AI's waifu-sona?)
Naturally does better with what I imagine are in-distribution[0] doodles/words than out-of-distribution[1], but still very cool (and fun!).
[0] https://scribblediffusion.com/scribbles/h373dd42xbduzerzhlos...
[1] https://scribblediffusion.com/scribbles/23xdfz5mtffwfgib32t7...
https://scribblediffusion.com/scribbles/nca3ivborbebhipju6jg...
Didn't get the squirrel part quite right, but maybe it's my drawing :-)
https://scribblediffusion.com/scribbles/bxc3jaofkzdh5nyjlff5...
But IMHO the interesting thing about controlnet is being able to use pre-rendered production art as a basis and allowing SD to respect the original proportions/model/etc. For rough sketches I prefer img2img without controlnet as it gives the algorithm leeway to fix, reinterpret, or "be inspired by" my input image without being too attached to it (since it's full of imperfections anyway).
I have no idea how that (have to spend hours just getting a prompt that looks semi decent) can be considered state of the art when something like midjourney (any prompt will produce something great) exists.
I feel like stable diffusion is mocking my scribbling skills
(using firefox hence the very poor drawing)
https://scribblediffusion.com/scribbles/i3p6uxpzf5cmbeety5nw...
https://scribblediffusion.com/scribbles/dppnzjesq5h5dfhjdadz...
I am disturbed by this.
Ie. Let me stage a mock up. Because I can’t draw a cat.
Practically, they are combined. (Triangle) + 'ice cream' = cone ; (Triangle) + 'mountains' = peak ; (Circle) + 'ice cream' = scoop ; (Circle) + 'mountains' = sun
Is the graphic used as a graphic?
In the brain, visual processing, left hemisphere was found to contain details; a right hemisphere to contain structural relations. So a whole is composed of elements and relative positions.
In Convolutional Neural Networks, "near, direct" layers contain analytic detail and "far, abstract" layers contain synthetic shapes.
So, implementation-wise, you can take e.g. descriptions as abstracts and a "pre-acquired" memory of details as «graphic».
Edit:
About the "combination", well that the whole purpose of this new technology proposal,
"ControlNet"
- i.e., formerly you may have had some "transformer" from input to output, and now "conditional controls" are added (through a "zero-convolution" technique) - see Adding Conditional Control to Text-to-Image Diffusion Models - https://arxiv.org/abs/2302.05543
“a photo of a cat”: https://scribblediffusion.com/scribbles/5ymr7r66yjgx3jaibns2...
vs:
“a photo of a cat”: https://scribblediffusion.com/scribbles/2noqq6amlzcpvpfsyweo...
How much trouble do you imagine registering a domain is?