llama.cpp: Roadmap May 2023
github.com
github.com
It should run on whatever as long as you have enough memory. How much exactly depends on the quantization mode chosen (it's a quality-memory-speed tradeoff), but you should expect to need between 0.5 and 1GB of memory per 1B parameters in the model.
Kobold cpp uses llama.cpp and provides a minimalist web UI. So if you want a chat assistant using llama.cpp on CPU, kobold cpp is probably what you want.