On the other hand, I bought an 8C/16T Zen 4 laptop with 64GB RAM and an 4TB SSD for less than $2000 total including tax. I’ll take that trade.
With a quantization of it you can run larger contexts and go a bit faster. 1.4 tok/sec at 8b quant with offload to a 6GB laptop GPU.
Speculative decoding has been being added to lots of the runtimes recently and can give a 20-30% boost with a 1 billion weight model running the speculative token stream.