For now I will just wait for AMD or Intel to release a x86 platform with 256G of unified memory, which would allow me to run larger models and stick to Linux as the inference platform.
Unfortunately, AMD dumped a great device with unfinished software stack, and the community is rolling with it, compared to the DGX Spark, which I think is more cluster friendly.
At 80B, you could do 2 A6000s.
What device is 128gb?
[1]: https://www.jeffgeerling.com/blog/2025/increasing-vram-alloc...
[1]: https://community.frame.work/t/ai-9-hx-370-vs-ai-max-395/736...
[2]: https://community.frame.work/t/tracking-will-the-ai-max-395-...
If you're targeting end user devices then a more reasonable target is 20GB VRAM since there are quite a lot of gpu/ram/APU combinations in that range. (orders of magnitude more than 128GB).