I fear that by the time the RTX Spark comes out it'd have to be $6k, and by the time a 128gb or more machine with 700+ GB/s comes out it'd be at $10k, way out of most consumers' hands.
Edit: capitalized gb/s to GB/s.
I fear that by the time the RTX Spark comes out it'd have to be $6k, and by the time a 128gb or more machine with 700+ GB/s comes out it'd be at $10k, way out of most consumers' hands.
Edit: capitalized gb/s to GB/s.
Waiting for the market to be less insane is somewhat akin to waiting for the s&p500 to drop a decent amount so you can buy in.
lol this is so wrong it's funny - equities go up in price, commodity goods go down in price. the two markets are literally diametrically opposed.
So, you should get into RAM futures if you believe this is more than a transient arbitrage sort of situation. All extant RAM will become obsolete as the demand shifts to newer, fancier versions.
xAI effectively did and lucked out to cover their losses and more with it.
Yes but on what timescale? Replacing the > 5 year old sticks in my laptop would currently cost well over half what the machine ran me when it was brand new.
I said less insane, not sane. If prices go down 20% but are still up 150% over a few years ago, that would still be an improvement from now.
Or it could work the other way: if new hardware comes out in 2027 such that the tokens/$ ratio works out better, that would also be less insane.
There will be a basic M6.
There was a post recently, that showed the DGX spark is twice as fast as the M3 Ultra at prompt processing, but half as fast at output tokens [2]. They used gps-oss-120 for that test with a small context.
[1] https://gpuquicklist.com/apus?models=GB10%20Grace%20Blackwel...
[2] https://aimultiple.com/dgx-spark-alternatives , https://news.ycombinator.com/item?id=48732679
I mostly run them clustered for DS4 and am quite happy with the performance, and the cost isn’t that much more for two than the MBP while giving me double the unified memory.
I’ll probably pick up a third to run multiple smaller models. I don’t understand why people would buy a halo over a spark at comparable prices, particularly because if you want to cluster, the cx7 be beats the shit out of them when it comes to latency and throughput
I like my Strix Halo and keep it chewing on stuff, mostly non-interactive workloads (security audits of software mostly, training experiments, etc.), I get a lot of use out of it. If you want to experiment with AI, it is a good platform for that, though at $4k you can get an Nvidia-based Asus Ascend GX10, which is probably better. But, if you want a local model for interactive agentic use, you're going to be running either Qwen 3.6 or Gemma 4, which will fit comfortably on 2x64GB GPUs (even old GPUs will run them faster than the Strix Halo...I have dual Radeon Pro V620s which are faster, and they're six years old), or snugly on 32GB. A 48GB or 64GB Mac would run them well. Two Radeon AI Pro R9700 GPUs is probably the sweet spot, right now for GPUs. Not the cost of a good used car, like a 5090 or 4090, but plenty of memory and performance for local inference. Also, not finicky and weird and needing custom 3D printed fan shrouds like the old server GPUs on eBay.
At the moment, there just isn't a model that works better on a 128GB inference machine like this that don't also work fine on 64GB machines, which may be faster (very few 32GB GPUs will be slower, though I wouldn't recommend buying any GPU that isn't currently actively supported by the vendor drivers and CUDA or ROCm...so probably don't buy an MI50 or V100 or whatever).
ctx_tokens,prefill_tokens,prefill_tps,gen_tokens,gen_tps,kvcache_bytes 2048,2048,202.02,128,15.31,52184460 4096,2048,211.03,128,14.64,80373132 6144,2048,208.04,128,14.59,108561804 8192,2048,200.78,128,14.43,136750476 10240,2048,203.04,128,14.37,164939148 12288,2048,200.82,128,14.27,193127820 14336,2048,198.62,128,14.22,221316492 16384,2048,196.14,128,14.20,249505164 18432,2048,189.48,128,14.13,277693836 20480,2048,186.59,128,14.06,305882508 22528,2048,183.88,128,13.99,334071180 24576,2048,183.38,128,13.92,362259852 26624,2048,181.57,128,13.87,390448524 28672,2048,183.46,128,13.80,418637196 30720,2048,181.80,128,13.73,446825868 32768,2048,175.93,128,13.55,475014540 34816,2048,175.42,128,13.46,503203212
Which almost exactly matches the benchmark you just linked. Looks like it's possible to goose it to 15 tokens per second with a tiny context, but why would I want a giant model with a 2k context? DeepSeek is too big to be fast enough on a Strix Halo.
To be clear, if you think that's comfortable for interactive use, more power to you. But, I'm not waiting for that. I'll pay DeepSeek to host it for me. Their token prices are quite cheap and their cached tokens are even cheaper...and they have the most effective caching in the business, as far as I can tell. Even naively using the API you get 80-90% cached token rate. If you use Reasonix, you get ~98% cached token rate. I just built a feature for an app I'm working on for $0.10 for 20 minutes of work. Not bad at all.
I think once someone comes up with a machine which has both it will easily sell for $10000 and people will be queueing to buy it.