Please compare the same things: carrots VS carrots, not apples VS eggs.
Please compare the same things: carrots VS carrots, not apples VS eggs.
200B is probably a rough estimate of Q4 + some space for context.
The Spark has 4x the VRAM of a 5090. That's all you need to know from a "how big can it go" perspective.
With 128 GB of unified system memory, developers can experiment, fine-tune, or inference models of up to 200B parameters. Plus, NVIDIA ConnectX™ networking can connect two NVIDIA DGX Spark supercomputers to enable inference on models up to 405B parameters.You can do it, if you quantize to FP4 — and Nvidia's special variant of FP4, NVFP4, isn't too bad (and it's optimized on Blackwell). Some models are even trained at FP4 these days, like the gpt-oss models. But gigabytes are gigabytes, and you can't squeeze 400GB of FP16 weights into only 128GB (or 256GB) of space.
The datasheet is telling you the truth: you can fit a 200B model. But it's not saying you can do that at FP16 — because you can't. You can only do it at FP4.
If the 200B model was at FP16, marketing could've turned around and claimed the DGX Spark could handle a 400B model (with an 8-bit quant) or a 800B model at some 4-bit quant.
Why would marketing leave such low-hanging fruit on the tree?
They wouldn't.