Pretty much any modern machine will do you fine for ~5 tokens/s on a 7B (small end) model like Mistral-7B or llama2-chat-7B (or any of their respective fine-tunes). The computer you already have can probably do this.
It can run on CPU cores if you have enough RAM for the model. It feels like you have time warped back to the early 1990's and are talking to someone on a BBS (AKA the words appear slowly), but it is entirely functional if you have 32+GB of RAM.
For reference: I run 7B 5bit quantized models on a ryzen 7 5700G with 64Gb at 8 tokens/second CPU only.
It's not close to what you can get with a high end graphics card but for every day use it is alright and has headroom for bigger models. Upgrading CPU, RAM and Mainboard in a 10+year old PC cost me just 400€
Do you specify the 5 bit at runtime or it’s a certain download?
I haven't tested, but I suspect you might be able to get good CPU inference on this mini pc: https://www.aliexpress.com/item/1005005825981362.html. It costs ~$800 in the maxed configuration with 64GB RAM clocked at 5600Mhz
Hopefully, AMD wakes up and allows ROCm / HIP on their 7040 series, that would be a killer feature.
And AFAIK the Ryzen pro do have an "AI coprocessor" but unsure of the API to use them.