People keep repeating this but how does higher bandwidth (probably not 8x higher though) compensate for a lower amount of RAM?
It's not quite as silly as the people saying that 8GB in the base config 'feels' much faster than 8GB on a PC cause the drive/swap are "so fast" but still..
With consumer desktop CPUs having 2 "channels" (~102.4GiB/s).
And prosumer desktop/workstation CPUs (e.g. Threadripper) having 4 "channels" (~204.8GiB/s).
While apple is claiming 819.2GiB/s.
That is _max throughput is 4x/8x more_ (depending if you compare it to prosumer (fair comparison) or consumer (unfair comparison for a $7k system) hardware.
Now the _max_ part is important, Apple mainly reaches this by having more bandwidth, i.e. parallelism.
Mainly (oversimplified) the per "CPU memory channel" bandwidth for x86 is 64bit (for 51.2GiB/s) so the M1 Ultra is roughly comparable to having 16 "CPU memory channels" instead of 8 (prosumer) or 4 (consumer).
Some applications can take advantage of this nicely and will scale potentially even to 4x/8x speed, most probably will not some might even have neglible improvements. But applications which use a lot of RAM (as much as they can get) and most of the RAM they use is "warm" (e.g. doesn't just lie around with little access but isn't super "hot", i.e. highly contented either) will profit quite a lot.
On the other hand applications which use little RAM but the same small RAM region very heavily contented likely will hardly at all profit and will likely run faster on overclocked RAM on consumer systems.
Luckily for apple most the typical use case they sell their pro desktop model for belong more/mostly in the first category.
Additionally if you can run Linux on this system they might become _very_ interesting for some scientific applications for some users. I mean even e.g. Zen 4 EPYC CPU only have 12 "CPU memory channels" not 16 and it's much easier to put a desktop box "somewhere" then a server unit.
Side note: I say "CPU channel" in quotes because while it tends to be the marketing term things are more complicated in practice, e.g. in general DDR5 is splitting the 64bit channels into 32bit sub-channels, and just listing the channel width and throughput is still not painting the whole picture at all, e.g. the latency also matters for some applications (hence why OC can make sense) etc.
EDIT: Correction: The Threadripper PRO models have 8 "CPU memory channels" so it's just 2x on a fair comparison.
I do hope folks other than Apple can get fast! The new CAMM dual channel module will hopefully help reduce footprint but there's plenty of soldered down systems so it's not really required. It also surprises me there's not a GDDR based APU, except I guess too many games need both a bunch of system ram and video ram and 32GB vram is expensive and has significant lower draw. Apple going wide is really the obvious move, & doing it on package was the best way to do it, it seems.
AFIK Apple currently doesn't have a foodhold on gaming outside of phone/tablet games (where they are strong, but the games are used to LPDDR perf.).
And while many of the Graphic applications people do use would profit from it I'm not sure it's that big of benefit.
In the end a lot of users just do daily tasks on Apple laptops and for that battery life absolutely trumps any speed benefits GDDR gives.
I guess the desktop versions could/should use GDDR, but there are two issue: 1. Heat, 2. more design/variance in the CPU production supply chain.
The 2nd can drive up cost by quite a surprising high amount.
The 1st point I'm not supper sure about. But AFIK one big problem with things which are stacked on-die is heat as all the heat of the CPU needs to go through whatever is stacked onto it to reach the cooler. And while there are probably all kinds of tricks to improve this I could imagine that using a GDDR which has a higher power draw and produces more heat itself could make this more of an issue. But that is purely speculative.
I guess we have to see until around 2025+ to know if such stacked chips are more prone to die a early (i.e. <5 years) heat death. Especially some of the Air models without heat pipes could be at risk if used in a less climate controlled environment. Or it could be all perfectly fine. I'm looking forward to finding out.
AFIK it's not that e.g. AMD couldn't do that and use on-die LPDDR5, I mean that is a bit different but not sustainable harder then using their chiplet + X3D tech.
The problem is that for AMD it's not a good business decision. First they need to be price competitive (a problem Apple doesn't have). Most applications outside some areas scale much more with the speed/latency of RAM then with bandwidth increase (beyond some basic level). In turn for a lot of e.g. desktop Ryzen processor use cases which are not media processing there is not enough value into adding many more channels. Especially if you consider that for large parts of the prosumer/server space IT admins will be really unhappy with on-die RAM Apple can afford forcing it, AMD can not. The reason is that while for desktop systems RAM death is rarely an issue in the server/heavily used workstation spectrum RAM death is not uncommon. Even for media processing or multi VM servers going beyond a certain number of channels (less then 16) is unlikely to be a good financial decision. I would go as far as arguing if the M2 Ultra wouldn't be based on tightly gluing 2 processors together and that processors have 8 channels because they are sold with a focus for media processing it wouldn't have anywhere close to 16 channels (but for Apple the cost of having less then 16 channels with their design is higher then the cost of having them, partially because they also don't sell server they don't want to compete with accidentally and similar).
Where I'm going with this is that outside of some media processing targeted products they don't need to "get their throughput up to more competive level" and getting it to having a some additional 8 channel choices (e.g. for some high end laptops) would likely be good enough in practice. And in turn we are unlikely to see more. And in turn they have no reason to try tricks like using GDDR memory.
The way Apple can push hardware vendors here is less because of a need for many more channels, but because for a want.
To spitball some figures, the rx7600 is 13b transistors, 165watts, and runs off 128-bit 288GBps GDDR6 memory. Or take 2017's rx580, but which was 256-bit GDDR5 good for 224GBps, at similar power.
An APU is going to be considerably lower power than either of these discrete gpus. I think I somewhat overestimated AMD's ability to scale up their APUs to be throughput limited, in most cases. The 256GBps LPDDR5X memory they're planning should offer a nice bump over where we are be pretty good.
I'm a bit surprised to see the memory bandwidth not being as constraining a factor as I had first guessed. It's also seemingly bizarre how overbuilt it makes Apple's memory setup look.
It's just that most people don't do that on a daily usage basis.
Maybe not a I need 800GiB/s constraining factor but definitely a I don't want just 200GiB/s constraining factor AFIK.
The PS5 and XboxX have an AMD APU (CPU+IGPU) with a wide memory interface. Seems like a fine decision. What surprises me is they haven't brought it to low/medium range desktops, until Strix Halo in 2024.
I guess it depends on your use case, but back when part of my day job was debugging performance problems with JVM-hosted applications, one of the things that was most noticeable was that the degree to which latency blows up memory use - whether the latency was GC, disk, network, DB queries, whatever. It all turns into holding items in memory longer before they get processed (which in the case of older JVMs, turns into a death spiral of GC thrashing, which blows up processing times further, until your application is staggering along).
Increasing memory can alleviate this, but only up to a point - and it can make things worse, because you've now got significant overhead managing the in-flight workloads.
It's also possibly that this is simply a decision driven by what Apple can produce with their M2 chips at this point, and that they wanted to offer 384 or 640 GB as the maximum, and everyone is making excuses for them.
Cause, I just added the GPU and system RAM bandwidth numbers together. Which is what needs to be kept in mind with much of this. Yes that is a lot of memory bandwidth and its hella useful for some subset of users, but its shared, and largely pointless for a lot of CPU bound tasks. But OTOH, may not be enough for many GPU bound ones.
It also assumes that pretty much every other CPU manufacture on the planet are idiots for optimizing for latency and putting in large caches to compensate (aka the desktop parts from AMD/intel have only _two_ channels, vs the 8+ in the server/workstation parts) and price discriminating for the parts that have more CPU bandwidth. AKA, you can get amd machines in the same ballpark (or possibly faster depending on how fast you can get 24 channels of DDR5 to run).
So, I'm not saying which is better because its likely workload dependent, but to claim its a blanket insurmountable advantage is questionable. Particularly since the price ranges we are talking about a similar machine is probably a 64 core threadripper plus a fat nvidia GPU or four and the shear core count and raw GPU compute is probably a win in most workloads.
Or if a GPU code needs more than 12-16GB of memory (normal cards) or 24GB (if you get a 4090)?
What I like about the apple approach is that low end laptops/desktops get 100GB/sec. Pay another $500 get 200GB/sec. Pay another $500 get 400GB/sec. Pay another $1000 get 800GB/sec and still fits in a small desktop. On the PC side with AMD/Intel you get the same memory bandwidth for the low, medium, and high end chips. Until you upgrade to a threadripper, which is a 280 watt chip, on an expensive motherboard, usually in a rather large PC case and makes the mac studio look cheap.
https://www.anandtech.com/show/17024/apple-m1-max-performanc...
(I don't know how well AMD's current processors do with utilizing the socket's full DRAM bandwidth from a limited number of chiplets, but I wouldn't be surprised if it's a more severe limitation than what M1 Max/Ultra show with their CPU clusters. It looks like only the 12-chiplet EPYC processors actually use all the links from the IO die to the CPU chiplets.)
So the inability to use all the DRAM bandwidth from the CPU cores, while perhaps disappointing, isn't exactly a weakness for Apple's processors compared to the competition.
Thinks like McCalpin do not seem to show much difference on the different number of chiplet Epycs, although I've not personally tested the newest Genoa chips.
Two data points from Apple:
a) M2 Ultra 24 cores, 128GB ram, 2TB storage = $5,200
b) M2 ultra 24 cores, 192GB ram, 4GB storage = $6,600
Likely 1/4th the size, 1/4th the power consumption, and 4x the ram bandwidth. Have you by chance played with any LLMs? Just saw a post that someone managed 5 tokens/sec with the llama 65B model.