My 3060 can generate a 256x256 8 step image in 0.5 seconds, no A100 needed. A 3090 is double the performance of a 3060 at 512x512, and an A100 is 50% faster than a 3090...
If you have access to an high end consumer GPU (4090) you can generate 512x512 images in less than a second, it's reached the point that you can increase the batch size and have it show 2-4 images per prompt without adversely affecting your workflow.
Too bad SD1.5 is too small* and we'll require models with more parameters if we want a true general purpose image model. If SD1.5 was the end-game, we'd have truly instant high res image generation in just a couple more generations of GPUs, think generating images in real time as you type the prompt, or have sliders that affect the strength of certain tokens and see the effects in real time, etc. Tho I heard that SDXL is actually faster for higher resolutions (>1024x1024) due to removing attention on the first layer, making it scale better with resolution even tho SDXL has 4x the parameter size.
* Current SD1.5 models that can generate consistent high quality images have been fine-tuned and merged so many times that a lot of general knowledge has been lost, e.g. they can be great at generating landscapes, but lacking in generating humans, or they can be very good at a certain style like comics but can do comic style only and lose the ability to generate more dynamic face variations, etc.