I feel like x86 offerings often fall down so flat in this area. Even when theoretical memory bandwidth is provided, there are often other significant limitations that prevent its use. I've run into this personally with AMD Zen cores, where it is pretty easy to saturate the infinite fabric bandwidth.
Do you know why they would outfit it with LPDDR5 instead of regular DDR5? I don't have a good feel for the difference here. At first blush I would assume low power DDR would run slower not faster, but that doesn't seem to be the case anywhere I see it deployed.
You can see line fill buffer issues with 8-socket machines today, for example, where they are the cause of stalls due to the high latency of cache coherency checks. However, since they are part of the core and fine for 99% of users, the size of the line fill buffer is kept relatively small.
That limitation puts the HBM Xeon Max sort of on par with Grace for actual usable memory bandwidth, not so far above it.
He talks about cache misses as 60ns, which glosses over that approximately half of that is missing through L1/L2/L3, then you enter the queue for the memory controller, for the memory channel you need. As a result you only get half the bandwidth if you only have a single request pending per channel.