https://arstechnica.com/gadgets/2025/08/review-framework-des...
https://arstechnica.com/gadgets/2025/08/review-framework-des...
To me this reads like "if you can afford those 256GB VRAM GPUs, you don't need PCIe bandwidth!"
That's pretty small.
Even Deepseek R1 0528 685b only has like ~16GB of attention weights. Kimi K2 with 1T parameters has 6168951472 attention params, which means ~12GB.
It's pretty easy to do prompt processing for massive models like Deepseek R1, Kimi K2, or Qwen 3 235b with only a single Nvidia 3090 gpu. Just do --n-cpu-moe 99 in llama.cpp or something similar.
[1]: It sounds like a nitpick but a PCIe x16 with x4 effective bandwidth can exist and is a different thing: if the actual PCIe interface is x16, but there is an upstream bottleneck (e.g. aggregate bandwidth from chipset to CPU is not enough to handle all peripherals at once at full rate.)