Perhaps they are referring to default GPU allocation that is 75% of the unified memory, but it is trivial to increase it.
Using an M3 Ultra I think the performance is pretty remarkable for inference and concerns about prompt processing being slow in particular are greatly exaggerated.
Maybe the advantage of the DGX Spark will be for training or fine tuning.