But a portable (no install) way to run llama.cpp on intel GPUs is really cool.
But a portable (no install) way to run llama.cpp on intel GPUs is really cool.
Requirements:
380GB CPU Memory
1-8 ARC A770
500GB DiskI think changing the end of headline to "Xeon w/380GB RAM" would stop it from being incorrect and misleading.
Edit: but what you added in your edit is right, it would be more accurate to append the system ram requirement
More GPUs let you keep more experts active at a time.
You might still be right since I have not confirmed that the selected experts change infrequently doing prompt processing / token generation, and someone could have botched the headline. However, treating Deepseek like llama 3 when reasoning about VRAM requirements is not necessarily correct.
[1] https://www.amazon.com/NEMIX-RAM-DDR4-2666MHz-PC4-21300-Redu...
The reason they used a Xeon is memory channels. Non-server CPUs only have 2 but modern Xeons have 8 to 12 depending on generation/type. And the Xeons with the most are the most $$$$ and it ends up cheaper to just get a GPU or dedicated accelerator.
It's a bit less exciting when you see they're just talking about offloading parts from the large amount of DRAM.
Also, LM studio lets you run smaller models in front of larger ones, so I could see having a few GPU in front really speeding up using R1 for inference.
Ktransformers has a document about using CPU + a single 4090D to reach decent tokens/s but I'm not sure how much of the perf is due to the 4090D vs other optimizations/changes for the CPU side https://github.com/kvcache-ai/ktransformers/blob/main/doc/en... The final step of going to 6 experts instead of 8 feels like cheating (not a lossless optimization).
K (K=8 for these models, but you can customize that if you want) experts of 256 per layer are activated at a time. The 256 comes from the model file, it's just how many they chose to build it with. In these models there is also 1 shared expert which is always active in the layer. The router picks which k routed experts to use each forward pass and then a gating mechanism combines the outputs. If you sum the 1 shared expert + K routed experts + router + output networks you end up with 37 B parameters active for each feed forward layer pass. The individual experts are therefore much smaller than the total (probably something like 4 B parameters each? I've never really checked that directly).
Or, for the short answer: "37 B is the active parameters of 9 experts + 'overhead', not the parameters of a single expert".
Two GPUs or more mean you can start to "keep" one or more of the experts hot on a GPU as well.
DeepSeek employs multi-token prediction which enables self-speculative decoding without needing to employ a separate draft model. Or at least that's what I understood the value of multi-token prediction to be.