Running Stable Diffusion in 260MB of RAM
github.com
github.com
A raspberry pi zero 2W seems to use about 6W under load (source: https://www.cnx-software.com/2021/12/09/raspberry-pi-zero-2-... )
So if it takes 3 hours to generate one picture, that's about 18Wh per image.
A Nvidia Tesla or RTX GPU can generate a similar picture very quickly. Assuming one second per image and 350W under load for the whole system it's in the magnitude of 0.1Wh per image.
Of course we could consider that a raspberry pi zero uses a lot less ressources and energy to be manufactured and transported.
And now I think I know what my next project is going to be, I am sure I can find some desk space
Edit..: I'm so hyped about this; the example image on TFA takes +2 hours to generate, but who cares?! I'd love to have a little display that churns around in the background and creates a new variation on my prompt every whatever hours, displaying the results on an unobtrusive eink screen.
but on the other hand I would also love the statement behind something unconnected to the internet that's slowly churning out unique, ephemeral pictures. Yours to enjoy, then gone forever.
[1] https://imgur.io/a/NoTr8XX (no, I don’t know why anyone would use Imgur to write up a hack either)
I'm thinking of making a Stable Diffusion version of this, and preferably with a larger eInk screen.
https://www.stavros.io/posts/making-the-timeframe/
You just put an image on some HTTP server and it shows it.
Check back in a few months for my results...
update: "Tests were run on my development machine: Windows Server 2019, 16GB RAM, 8750H cpu (AVX2), 970 EVO Plus SSD, 8 virtual cores on VMWare."
I wonder how fast would a consumer PC, with no GPU, generate an image with say 16gb of RAM?
[1] https://github.com/danieldk/gemm-benchmark#example-results
What is the resolution of your images and number of steps?
Edit: confirmed.
> Prompt length shouldn't influence creation time...
Yeah, checks out with my experience too. Longer prompts were truncated.
Albeit in 77 token increments.
DDIM, 12 steps.
Apparently there are ways around it, but I just switched to runpod.io. It's very cheap (around $0.80/h for a 4090 including storage) and having a real terminal is worth it.
While true, neither your statement or mine above is germane to the discussion. It wasn't about how long it takes. It's a discussion of how cool it is that it can be done on that machine at all.
If the random number generator is pseudo-random, this makes GPT-4 a deterministic finite-state machine, and the output sequence does not necessarily contain all possible subsequences no matter how many times the monkey types a new random key. Put differently, some output subsequences may be inaccessible no matter which keys are input. Same if the random number generator is truly random but its value cannot select among all possible output tokens, only a subset provided by the GPT at each step.