Mostly I'm interested in the processing time. Like, using a midrange desktop, what's the average time to expect SD to produce an image from a prompt? Minutes/Tens of minutes/Hours?
Mostly I'm interested in the processing time. Like, using a midrange desktop, what's the average time to expect SD to produce an image from a prompt? Minutes/Tens of minutes/Hours?
Keep in mind to have the batch-size low (equal to 1, probably), that was my main issue when I first installed this.
Then, there's lot's of great forks already which add an interactive repl or web ui [0][1]. They also run with half-precision which saves a few bytes. Additionally, they optionally integrate with upscaling neural networks, which means you can generate 512x512 images with stable diffusion and then scale them up to 1024x1024 easily. Moreover, they optionally integrate with face-fixing neural networks, which can also drastically improve the quality of images.
There's also this ultra-optimized repo, but it's a fair bit slower [2].
[0]: https://github.com/lstein/stable-diffusion
(I say “I think” because I’ve uninstalled the nvidia-dkms package again while I’m not using it because having a functional NVIDIA dual-GPU system in Linux is apparently too annoying: Alacritty takes a few seconds to start because it blocks on spinning up the dGPU for a bit for some reason even though it doesn’t use it, wake from sleep takes five or ten seconds instead of under one second, Firefox glyph and icon caches for individual windows occasionally (mostly on wake) get blatted (that’s actually mildly concerning, though so long as the memory corruption is only in GPU memory it’s probably OK), and if the nvidia modules are loaded at boot time Sway requires --unsupported-gpu and my backlight brightness keys break because the device changes in the /sys tree and I end up with an 0644 root:root brightness file instead of the usual 0664 root:video, and I can’t be bothered figuring it out or arranging a setuid wrapper or whatever. Yeah, now I’m remembering why I would have preferred a single-GPU laptop, to say nothing of the added expense of a major component that had gone completely unused until this week. But no one sells what I wanted without a dedicated GPU for some reason.)
With an RTX 3060, your average image generation time is going to be around 7-11 seconds if I recall correctly. This swings wildly based on how you adjust different settings, but I doubt you'll ever require more than 70 seconds to generate an image.
I quickly threw together a folder structure where I have a md5'd prompt as a folder name, into that goes _promp.txt with the actual text of the prompt and the images i generate in a loop with the seed used and iteration number in the image's file name. That way I can generate like 20 seed-based images for a prompt and if the model bites I let it run with a much higher number of seed-based images. When you have a 1000 to pick from, some of the results are freaking amazing.
My first impression is it seems a lot more useful then DALL-E, because you can quickly iterate on prompts, and also generate many batches, picking the best ones. To get something that's actually usable, you'll have to tinker around a bit and give it a few tries. With DALL-E, feedback is slower, and there's reluctance to just hammer prompts because of credits.
I was able to obtain 256x512 images with this card using the standard model, but ran into OOM issues.
I don't mind waiting, so now I am using the "fast" repo:
https://github.com/basujindal/stable-diffusion
With this, it takes 30s to generate a 768x512 image (any larger and I am experiencing OOM issues again). I think you should expect a bit faster at the same resolution with your 3060 because it's a faster card with the same amount of memory.
It runs pretty well but the most I can get is a 768x512 image, but it's pretty good for stuff like visual novel background art[0] and similar things.
[0] - https://twitter.com/xMorgawr/status/1564271156462440448