Image generation takes 5 sec, but loading the model takes 120 sec. There's some time you can keep the model "alive" for, to wait for another image generation request, but if I have too few users, it means everyone gets image generation times of 120 sec and I pay for 120 sec per image, which is a lot.
With enough users, though, I can always keep the model up, and everyone gets 5 sec generation times and the economics for this make sense. Hopefully if this picks up steam I'll be able to rent a dedicated server to generate things on, but it's just a sideproject for now.
At least I hope it will be useful/fun to some people.
EDIT: Added, good shout, thanks!
A few days ago I was generating images using my home desktop, and had a Web push notification notifying people when it was on. That was much cheaper, but very bad UX, as it could be hours before you got your image (and you were long gone by then).
However, I can see why you might choose not to bother with that and find another way to keep the model alive without the try prompt.
Edit: didn't you mention it takes 5 sec if model is alive?
Also, Inferrd doesn't really look cheaper, a 16GB GPU is $600/mo, which is about the same as Banana, except Banana doesn't bill for the time the model isn't working.
For my needs, with the few images I need to generate, the difference is between paying $50/mo (admittedly, with worse UX for my customers) and paying $600/mo for a card that's mostly sitting idle.