Can any of this realistically run on CPU at some point?
(Not training obviously)
(Not training obviously)
There a lot of optimizations that can be done. Here's one w/ potentially a 15X AVX speedup for example: https://github.com/ggerganov/llama.cpp/pull/996
see this comparison: https://old.reddit.com/r/LocalLLaMA/comments/12ezcly/compari...
these models quantised to 4bit should run in CPU set ups with 16GB of RAM + 16GB of swap (Linux) and perhaps other setups run similarly