My biggest problem with running local LLMs on my M4 Max/128GB RAM is the prefill latency.
I've since acquired two DGX Sparks, and it feels so much snappier.
I've since acquired two DGX Sparks, and it feels so much snappier.
the sparks have much slower memory bandwidth is the trade off
Another benefit of the 2x spark setup is that you can parallelize to ~6 streams pretty efficiently.
All depends on the workflows you’re using it for.
I’m quite excited for the M7 class machines.