Think about it, give the user a few basic, MS-Paint level pencil tools, colors, shape makers. Ask for a description, the application can even push you in the right direction for putting together good, detailed prompts, gives you a list of art-styles, artists, filtering methods, etc all with reference images so you don't need to memorize names. You can zoom into sections of the image to work on independently (like the birds in the article), then blend it into the greater image. Drag and drop image files onto the project and iterate on them.
Implementing the glue to simplify the "tough" parts of this process is honestly pretty trivial.
Using a CLI-based tool is inaccessible for most people... but building a GUI around this would be very easy. I'm too lazy to google it, but I would bet someone already has a GUI, or is working on one.
12GB of VRAM may not be accessible on most computers, but there's nothing innovative about offloading that task to an EC2 instance. It just requires an opportunistic developer to tie the pieces together.
I would be monumentally surprised if Figma/Canva/InVision/Adobe are not already working on this.
CLI-based tools are perfectly accessible to most people.
They just can't be arsed to learn them, unless they need to. And most of the time, they don't, because good-enough alternatives exist.
If a CLI-based tool is the only way that an average person can get their work done, that's what they'll use.
There's a WebUI with a docker container if you're on Linux w/ GPU; https://github.com/AbdBarho/stable-diffusion-webui and https://github.com/AbdBarho/stable-diffusion-webui-docker.
If you don't have a GPU, there's a Colab UI (Google hosted GPU). https://github.com/pinilpypinilpy/sd-webui-colab-simplified
Probably not many in general, but the RTX 3060 has 12GB or ram and it is around $350. And I saw a RTX 2060 12GB for $250 the other day. That's a pretty reasonable entry fee IMO.
Itll comfortably run on 6gb now. gtx1600 series cards need to run in full precision mode to produce output. The HLKY fork has improved the Gradio GUI and integrated realesrgan and gfpgan for those with beefier cards.
Someone else also figured out how to load and run it all on a CPU, so pretty much anyone can in theory run the model now.
There is an elaborate Colab notebook linked in the HLKY repo that seems to get more point and click user friendly every time i look at it. I think it even launches the gradio webui so you can use the Colab instance with a webui remotely.
The neat new applications that have taken over this site for the last couple days sometimes require CLI steps to install because they are in active development and it can be easier to experiment with something local. I'm sure they'll either be moved online or wrapped in nice installers over the next couple weeks.
4 out of 5 people globally would be able to submit a stable diffusion prompt and view a result. Most would have no idea what the hell was going on or even why it was interesting.
This is the funniest part to me, because so many people already think this is how digital art worked to begin with.
Yet. This is a huge leap forward, to get more basic prompts generating things will be a much smaller leap IMO.