The big advantage of the DGX Spark over the Strix Halo is much faster prefill. Like 5x the speed. Also the networking hardware on it is insanely powerful, though I and 99% of other Spark users, are unlikely to use it to its full capacity.
The big advantage of the DGX Spark over the Strix Halo is much faster prefill. Like 5x the speed. Also the networking hardware on it is insanely powerful, though I and 99% of other Spark users, are unlikely to use it to its full capacity.
Let's just say if I had my druthers, I would not choose Ubuntu, and I really wouldn't choose the Nvidia/ARM spin of Ubuntu. The Strix Halo has the benefit of being entirely a normal x86_64 PC that happens to also have a big chunk of unified memory. You can put pretty much anything on it. Any Linux, regular Windows, probably even a BSD (though good luck getting AI stuff working there).
But, as I said, if you're spending $4000 exclusively for inference and AI workloads, you might as well get the Nvidia-based unit. It is better for that.
Regular Ubuntu or Fedora ISOs do work out of the box
Just needs the Nvidia GPU driver install afterwards
(And the realtek 10gbe module oot, or blocklist if not used)
For Jetson Orin and later:
You can download an ISO from https://developer.nvidia.com/embedded/jetpack/downloads and that'll work. Other distributions are still a bit of a mess though but Yocto is supported now.
I basically treated it as a stock Ubuntu machine but on Aarch64, and ignored the stuff they installed. And then went looking for docker images for vllm that made sense and were up to date and hopefully tuned for NVFP4 on the hardware and...
... that was more the disappointing part.
Still, I'd rather have the blazing prefill speed of the Nvidia over my Strix Halo, but I wasn't willing to spend nearly twice as much for it (when I bought, the Strix Halo was $2100 and the GX10 was going for $3700, I think). Now that the difference between the two is much smaller, there's no reason to get the AMD.
I am not sure why NVIDIA couldn't have released a cheaper version of the thing with just standard Ethernet on it and leave it at that. Hardly anybody can afford two or more of them to cluster via ConnectX anyways.
I'm running vllm + GoModel on k3s, works like a charm. Even wrote some CUE last night to generate the values file for my Helm chart. It calculates the GPU percentage for me from the memGB I assign to a model.