To run the full 671B Q8 model relatively cheaply (around $6k), you can get a dual EPYC server with 768GB RAM - CPU inference only at around 6-8 tokens/sec. https://x.com/carrigmat/status/1884244369907278106
There are a lot of low quant ways to run in less RAM, but the quality will be worse. Also, running a distill is not the same thing as running the larger model, so unless you have access to an 8xGPU server with lots of VRAM (>$50k), cpu inference is probably your best bet today.
If the new M4 Ultra Macs have 256GB unified RAM as expected, then you may still need to connect 3 of them together via Thunderbolt 5 in order to have enough RAM to run the Q8 model. Assuming that the speed of that will be faster than the EPYC server, but will need to test empirically once that machine is released.