Opus 4.5 level of performance is also accessible with deepseek-v4-flash-0731 (0731 being the july 31 update) which is much, much, much smaller. 2x RTX pro 6000 blackwell can run it. 4x can run it very comfortably
I am running DS v4 flash 0731 lossless at 80t/s right now. It really is not at Opus 4.5 level (for my workload). I would say it's around 3.7 Sonnet, which is still pretty good, but other models such as GLM 5.2 are still leaps better. Of course I run DSv4 flash over GLM 5.2 for a few very good reasons, but intelligence is not 1 of them.
Despite fitting into VRAM, I can't get DSV4 to run at usable speeds on my AMD hardware. The upcoming qwen3.8 27b greatly excites me, and I hope it can outperform Stepfun 3.7 Flash, which is the best thing I can run today.
I'm just trying out Muse-Glimmer 30b, and my initial vibe is this might be better than qwen3.6-27b. No idea how it compares to Stepfun, because I can't run that model - but worth checking out while you wait for qwen3.8-27b
Anyone thinking of buying 2x RTX Pro 6000 Blackwells - beware: unlike other cards e.g. RTX 5090, The RTX Pro 6000 cards cannot be NV-Linked, so you'll be going through the PCIe bus instead (7x higher sync cost)
My understanding is the last consumer card that supported that was the 3090. A Google search seems to agree the 5090 does NOT support NVLink...
What do you need the extra 2 for? Tensor parallelism?
Longer context and more cache. The problem is that native format with DSpark enabled you have very little room on the VRAM.
I was under the impression that you could fit the full 1M context within the 192GB VRAM as a result of DeepSeek's various architectural advancements, but I'll grant that DSpark + a larger pool for concurrency may necessitate more VRAM, yes.