https://github.com/lucidrains/DALLE2-pytorch/discussions/10
Inference cost and scale seems to be much more favourable than large language models (for now).
TL;DR: A single NVidia A100 is most likely sufficient; with a lot of optimization and stepwise execution a single 3090 Ti might also be within the realm of possibility.
https://github.com/lucidrains/DALLE2-pytorch/issues/22#issue...
GPT was rumored to cost in the millions, perhaps the hours estimate is conservative?
(Small timer with absolutely no idea of practical ML here)
[1]: https://github.com/alembics/disco-diffusion
[2]: https://twitter.com/midjourney?t=-kKC5UE-gjIkMvAb709SyQ&s=09
The Latent_Diffusion_LAION_400M notebook generates six 512x512 images in about 45 seconds on a K80 on Collab.
DALL-E2 is more complicated but presumably also better optimised.