> If Apple M4 Ultra gets close to 1.8 TB/s bandwidth of 5090
If past trends hold (Ultra = 2x Max) it'll be around 1.1 TB/s, so closer to the 4090.
> If Apple M4 Ultra gets close to 1.8 TB/s bandwidth of 5090
If past trends hold (Ultra = 2x Max) it'll be around 1.1 TB/s, so closer to the 4090.
More tokens/sec on the dual 5090, but way bigger model on the M4.
Plus the dual 5090 might trip your breakers.
128GB of VRAM for $3000.
Slow? Yes. It isn't meant to compete with the datacenter chips, it's just a way to stop the embarrassment of being beaten at HPC workstations by apple, but it does the job.
32GB (for 5090) / 24GB (for 4090) ≃ 1.33
Then multiply 4090's price by that: $1500 × 1.33 ≃ $2000
All else equal, this means that price per GB of VRAM stayed the same. But in reality, other things improved too (like the bandwidth) which I appreciate.I just think that for home AI use, 32GB isn't that helpful. In my experience and especially for agents, models at 32B parameters just start to be useful. Below that, they're useful only for simple tasks.
a) they are never going to be happy,
b) it's actually a significant step up given the bulk are dual-card users anyway, so this bumps them from (at the high end of the consumer segment) 48GB to 64GB of VRAM, which _is_ pretty significant given the prevalence of larger models / quants in that space, and
c) vendors really don't care terribly much about the home / hobbyist LLM market, no matter how much people in that market wish otherwise.
Buy their datacenter GPU if you need more VRAM.