AMD’s Zen 4, Part 2: Memory Subsystem and Conclusion
chipsandcheese.com
chipsandcheese.com
Flight sims (especially DCS) tend to be built on archaic engines that rely heavily on non-parallel workloads, so the stacked cache approach can yield some crazy performance gains. I'm excited for the upcoming Zen 4 version of their X3D part.
768MB (!) of L3 cache. Looks like it's only Azure for the time being. Couldn't find those parts on AWS or GCP.
Although, AMD is using their own processors via GCP to.. design their own processors.[2] However, it doesn't seem like they're using 3D V-Cache parts there.
Exciting times in any case. There's a lot of workloads that stand to benefit from these new cache architectures.
[0] https://www.storagereview.com/news/amd-epyc-3rd-gen-processo...
[1] https://www.hpcwire.com/2022/03/21/amd-milan-x-cpu-with-3d-v...
[2] https://insidehpc.com/2022/05/amd-to-use-amd-epyc-powered-go...
For example, if you want to store a data, this data will first go into a store-buffer and core at this moment is basically done. This operation can go only up to 64 bytes per single cycle per single ALU port. Skylake had only one such port (Store Data) whereas Sunny Cove upgraded to having two such ports. In practice this means, provided that you have at least two store uOps in the CPU uOps pipeline (which maxes out at 4-5 uOps per cycle), Sunny Cove could double the bandwidth because it could store 128 bytes into a store buffer per each cycle. Buffering in general and regardless of the uarch is helping to hide the memory subsystem latencies. And I guess these could be the "bursts" that you might have read about.
End of game to hiding the latencies though is when your code demands that the data you want to store must be immediately visible to other cores. In that case you have to flush the store-buffer which, along with the pipeline flush due to branch misprediction, is one of the most expensive operations you can do in x86-64.
BTW, I'm pretty sure in sunny cove they just went from 1x 64-byte store port to 2x 32-byte store ports, so the actual bandwidth did not increase for vectorised code.
(technically, the way AMD did v-cache in zen3 meant that cache bandwidth didn't change, just higher hitrate, but, it doesn't mean it will always be done that way in the future. RDNA3 saw a shift from a focus on capacity/hitrate towards higher bandwidth - infinity cache stayed the same size but much higher bandwidth, which of course requires more transistors. Maybe we will see something similar on Zen4 - you could have a Crystal Well-style L4 or 6775R-style side-cache. Or even both a side-cache and a big L3 on the same design - just stack them.)
ARMs have a looser memory model and as a result often manage a much higher fraction of the peak memory bandwidth. Power as well.