Nvidia's latest AI PC boxes sound great – for data scientists with $3k to spare
theregister.com
theregister.com
You can buy a Mac Studio with an M4 Max for $3,500. 128GB unified memory @ 546 GB/s.
You're also getting a much faster CPU and more usable daily computer.
I suppose if you're a CUDA developer, this thing is probably better though I doubt you'd be training anything worthwhile on a computer this weak. Nvidia advertises the DGX Spark as a machine that mimics very large DGX clusters so the environment is the same. But in terms of hardware specs, it's very disappointing for $3,000.
DGX Station is another beast. It's Blackwell Ultra with 288GB HBM3e and 496GB LPDDR5X. I'm guessing $150k - $200k.
Time point 2 is replying.
Because the bottleneck to producing a single token is typically the time taken to get the weights into the FPU macs perform very well at producing additional tokens.
Producing the first token means processing the entire prompt first. With the prompt you don't need to process one token before moving on to the next because they are all given to you at once. That means loading the weights into the FPU onlu once for the entire prompt, rather than once for every token. That means the bottleneck isn't the time to get the weights to the FPU, it's the time taken to process the tokens.
Macs have comparatively low compute performance (M4 Max runs at about 1/4 the FP16 speed of the small nvidia box in this article, which itself is roughly 1/4 the speed of a 5090 GPU).
I have a M4 Max with 128 GB RAM. If I play Civ VI (a game released in 2016) on it for a few hours without limiting the FPS it will heat up till it turns itself off.
It's not cut out for heavy loads consistently like games or crypto mining. It's cut out for a heavy loads like Xcode compiling for a few minutes then back to editing text.
My gaming machine which has poorer performance (in fps terms) and is equipped with AMD 3700 and a 2080 Super can play Civ VI indefinitely without breaking a sweat.
If someone is looking at this Nvidia box, they're likely fine with a desktop footprint, in which case they'd be looking at the Mac Studio which should not have any thermal issues whatsoever. I'm guessing you're on a laptop?
If they are insistent on a laptop format, you can alleviate overheating issues pretty easily with some thermal pads and running the laptop on a cooling base when you're running heavy operations:
I've had no trouble with a (base) M4 mini for regular dev work, though I compile remotely and haven't played any games on it.
Shocking to say the least imo.
I think the best you can dream of is 480.0 GB/s, so 447 GiB/s.
TLDR, AMD machines based on HX395 offer comparable memory bandwidth and size at lower cost ($2000) but lack the high speed networking and compute power (60TF FP16 on AMD and 120TF on NVIDIA)
Apple M4 Max costs 23% more with equivalent memory, has twice the memory bandwidth but again no fast networking and significantly less compute (30TF FP16 on M4 Max).
For inference on large language models the extra memory bandwidth probably means M4 Max is the fastest option unless you use large batches or long contexts. For training large models needing more than 128GB RAM two of these Nvidia boxes is probably fastest.
[0] https://rocm.docs.amd.com/projects/install-on-windows/en/lat...
[1] https://rocm.docs.amd.com/projects/install-on-linux/en/lates...
I think people are running LLMs on HX395 using the Vulkan backend in llama.cpp which will work on Linux. No RocM or pytorch though.
>You also have the option to install the NVIDIA DGX Software Stack on a regular Ubuntu 22.04 while still benefiting from the advanced DGX features. This installation method supports more flexibility, such as custom partition schemes.
https://docs.nvidia.com/dgx/dgx-os-6-user-guide/introduction...
I don't fully understood why they are separate from gaming computers other than marketing, but I am very happy with my tensorbook. I got it when I was very young and impulsive I don't know if I would get it again, although that fact alone speaks to the computer's longevity.
So, they make products preloaded with configurations for each market segment. The customers pay for all the headaches this saves them of custom setups.
The actual price will be 2-3x this from some scalper who has a giant pile of them while people who want to do actual work with them pay through the nose.
That's assuming the connectors don't catch on fire, no cores are missing the way ROPs have been missing with 5090s, etc.
So would seem like they’re managing the allocations pretty carefully, or at least trying harder than they have before.
A middle tier could be something like an RTX 5090-like chip connected to 128GB - 256GB of GDDR7 RAM for $15k - $30k.
It would be nice dual use system.
These were high-end boards that worked fine in Linux when I bought them (as long as you were willing to periodically patch + recompile the kernel from a text console).
Anyway, once they have feature parity with Windows + current CUDA with a 100% open source stack going back a decade, then I'll consider them again. Until then, I'll happily buy an AMD that's 90% as fast for the same price.
Any org with less than several hundred thousand dollars to spare isn't anyone that NVIDIA cares about at all...
None of this speaks to the actual GPU performance. The only spec the DGX Spark webpage mentions is “1,000 FP4 TOPS”.