I suspect a big part of why stable diffusion managed to consume so much mindshare is that it can run on ordinary consumer hardware. On that point, I would be excited about an open-source RETRO (https://arxiv.org/pdf/2112.04426.pdf) model with comparable performance to GPT-3 that could run on consumer hardware with an NVMe SSD.
The biggest bloom I personally have run on cpu only in this fashion is 7B. It requires 4x7B of RAM plus some. On my hardware it tends to use all 32GB RAM and about ~4GB of storage during inference. At the moment I believe there is still a limitation of the smallest layer fitting in memory at once. This is why I haven't tried bigger bloom, but I believe there are ways to overcome it. Once this problem is resolved one should be able to use the same tech to use GPUs with less vram (like my 2070 with 8GB) for parts of larger models.
- You give us a dollar, we'll give you four quarters!
- People ask us how we make money. The answer is simple: _volume_.