Why can't this scale to run much larger models on CPU backed by flash with good access patterns?
I think they got something like 10 _seconds per token_ (not tokens per second).
EDIT: it was 25GB of ram and up to 20 seconds per token! https://github.com/JustVugg/colibri
As someone with a healthy amount of RAM, but just a 16GB GPU, I am wondering what kind of work I could queue up for overnight runs. I thought the best models were fully out of reach, but the 128GB CPU only test had a 1.8 tokens/second. While not speedy, you could probably do something with that given extensive coffee breaks. This speed simulator[0] demos what it looks like.
I could see where you could set up some coding task and let it churn all night.