They call it that but it's really LPDDR5, i.e. normal DRAM, using a wide memory bus. Which is the same thing servers do.
The base M3, with "GPU memory", has 100GB/s, which is less than even a cheap desktop PC with dual channel DDR5-6400. The M3 Pro has 150GB/s. By comparison a five year old Epyc system has 8 channels of DDR4-3200 with more than 200GB/s per socket. The M3 Max has 300-400GB/s. Current generation servers have 12 channels of DDR5-4800 with 460GB/s per socket, and support multi-socket systems.
The studio has 800GB/s, which is almost as much as the modern dual socket system (for about the same price), but it's not obvious it has enough compute resources to actually use that.
GPUs have a lot of memory bandwidth. For example, the RTX-4090 has just over 1000GB/s, so a 40GB model could get up to 25 tokens/second. Except that the RTX-4090 only has 24GB of memory, so a 40GB model doesn't fit in one and then you need two of them. For a 128GB model you'd need six of them. But they're each $2000, so that sucks.
Servers with a lot of memory channels have a decent amount of memory bandwidth, not as much as high-end GPUs but still several times more than desktop PCs, so the performance is kind of medium. Meanwhile they support copious amounts of cheap commodity RAM. There is no GPU, you just run it on a CPU with a lot of cores and memory channels.
And yes, of course it's not magic, and in principle there's no reason why a dedicated LLM-box with heaps of fast DDR5 couldn't cost less. But in practice, I'm not aware of any actual offerings in this space for comparable money that do not involve having to mess around with building things yourself. The beauty of Mac Studio is that you just plug it in, and it works.
Please list quantization for benchmarks. I'm assuming that's not the full model because that would need 256GB and I don't see a Studio model with that much memory, but q8 doubles performance and q4 quadruples it (with corresponding loss of quality).
> But in practice, I'm not aware of any actual offerings in this space for comparable money that do not involve having to mess around with building things yourself.
You can just buy a complete server from a vendor or eBay, but this costs more because they'll try to constrain you to a particular configuration that includes things you don't need, or overcharge for RAM etc. Which is basically the same thing Apple does.
Whereas you can buy the barebones machine and then put components in it, which takes like fifteen minutes but can save you a thousand bucks.
And the Max has half as many cores as the Ultra, implying it would be compute-bound too.
The speed is certainly not comparable to dedicated GPUs, but the power efficiency is ridiculous for a very usable speed and no hardware setup.
I have one, where I selected an M1 Ultra and 128G RAM to facilitate just this sort of thing. But in practice, I'm spending much more time using it to edit 4K video, and as a recording studio/to develop audio plugins on, and to livestream while doing these things.
Turns out it's good at these things, and since I have the LLAMA 70b language model at home and can run it directly unquantized (not at blinding speed, of course, but it'll run just fine), I'm naturally interested in learning how to fine tune it :)
I still wouldn't recommend it to someone just looking for a powerful desktop, just because $3K is way overpriced for what you get (non-replaceable 1Tb SSD is so Apple!). But it's certainly great if you already have it...