Or did we lose Stable Diffusion to SAAS also? Like we did on many of the LLMs which started of so promising as for self hosting goes
Or did we lose Stable Diffusion to SAAS also? Like we did on many of the LLMs which started of so promising as for self hosting goes
> In early, unoptimized inference tests on consumer hardware our largest SD3 model with 8B parameters fits into the 24GB VRAM of a RTX 4090 and takes 34 seconds to generate an image of resolution 1024x1024 when using 50 sampling steps. Additionally, there will be multiple variations of Stable Diffusion 3 during the initial release, ranging from 800m to 8B parameter models to further eliminate hardware barriers.
I'm not even sure what the use case is.
I'd rather wait 30 seconds and get a much higher quality image than some mediocre image in 1 second.
Hell, even if it took 5 minutes per image, and produced even better images, I would prefer that.
Some use cases might be generating profile pictures or banners for users or unique profile pictures for bots in online games. Discord, steam, social media, whatnot, you could just type what you want your profile picture to be and make it on the fly. They're small, aren't expected to be extremely high quality, and cheap enough.
Testing on https://fastsdxl.ai/ - "high quality profile picture of a cartoon cat holding a Bouquet of flowers"
To be clear it's not perfect, but this is a fairly complex prompt and I find the majority of seeds would be "good enough" for thumbnail profile pictures. I think we're almost there for "cheap good enough" usecases.
Places like steam, discord, etc you very rarely see profile pictures above that size.
It is just art. I think AI art shows what obsessed gadget makers for profit we have become culturally. We can't even figure out that the use case for art is hanging on the wall for decoration. For a conversation piece. A few will have their name become known and make it into galleries.
Infinite supply means the value tends towards zero.Good luck monetizing anything with those economic characteristics.
30 seconds is probably a decent sweetspot imo.
If you are getting what you want most of the time, then you are a better 'prompt engineer' than I am.
For example it could take an old video game say morrowind and it could in real time patch the graphics onto the video screen. Or people could look at a video of themselves and it would update the style similar to a snapchat filter.
But the difference is academic; progress is so fast that it is reasonable to expect all these models will be obsolete in a year or two.