https://github.com/MoonRide303/Fooocus-MRE
For base SD 1.5, I use Volta, because its fast: https://github.com/VoltaML/voltaML-fast-stable-diffusion/com...
Really good SD 1.5 image quality comes from gratuitous use of finetunes, LORAs, controlnet and other augmentations. So you can, say, trace a base image for structure, specify prompting in certain areas of the image and so on. InvokeAI is actually quite feature packed, and has lots of these augmentations hidden in the nodes UI, but Volta and other UIs also expose them more directly.
Still, even with it turned off, the quality is quite remarkable.
Ip adapter uses an image to guide denoising.
Fooocus and MJ take a prompt and expand it in a variety of ways (eg a language model or more simplistic text manipulation). The actual prompt that creates the conditioning is not what you typed in. That’s what I mean by prompt massaging
Generally the trade off is that any of the impressive finetuned models are far less generalizable then the default weights, but in practice this is not a big deal and the results can be a substantial improvement.
You’ll need more time and memory compared to Invoke or an Nvidia graphics card, but it’s not that bad: 1-2 s/it for an image in standard 512x768px quality, 14-20 s/it for an image in high 1024x1536px quality (Hires Fix).