Wouldn't the HBM-equipped Xeon Max be the bandwidth king of these tables?
You can see line fill buffer issues with 8-socket machines today, for example, where they are the cause of stalls due to the high latency of cache coherency checks. However, since they are part of the core and fine for 99% of users, the size of the line fill buffer is kept relatively small.
That limitation puts the HBM Xeon Max sort of on par with Grace for actual usable memory bandwidth, not so far above it.
He talks about cache misses as 60ns, which glosses over that approximately half of that is missing through L1/L2/L3, then you enter the queue for the memory controller, for the memory channel you need. As a result you only get half the bandwidth if you only have a single request pending per channel.