"has a big reorder buffer", I'm interpreting this as "one of the notable advantages of the M1 is it's ability to be more clever about processing instructions out of order to maximize resource utilization". Is that about right?
If you’re really interested in this it might be worth finding a copy of the Patterson and Hennessy book. It’s a big read and expensive (but older versions are on the internet archive [1]) and covers all these design issues in quite a lot of detail.
[1] https://archive.org/details/ComputerArchitectureAQuantitativ...
All of this is SDRAM, the S stands for synchronous and means that the memory is driven according to fixed timings in relation to a fixed bus clock. All LPDDR4X-4266 parts with the same timings perform exactly the same, whether they are soldered to the interposer or are 5 cm away on the board.
I'll go first.
Personally when talking about memory latency, I want to measure time to main memory, without TLB thrashing. Not that TLB latency isn't interesting, but it should be separately quantified.
I've written a few microbenchmarks in this area, but I get very similar numbers to the Anandtech R per RV prange. Which puts the latency to main memory at around 35ns looking like it might flatten out at 40ns. Yes, full TLB thrashing is up at 110ns or so, but that's not a usual use case. If it is, at least under linux, you can switch to 1GB pages if that's important to you.
R per RV prange numbers on the Ryzen 9 5950x is 65ns or so.
So sure the TLB worst case is higher latency on the M1, but then you have to figure out how big the TLB is, and how often that's an issue if you want to know the real world performance impact.
The i9-12900K with DDR5 gets 30ns on the R per RV prange, and 92ns on full random (tlb thrashing).
Even assuming the worst case TLB behavior the m1 max is 111ns and the 5950x is 79ns or 1.4x higher. On the Intel side is 111ns vs 92n, 1.2x higher.
Finding it hard to find any numbers that make the M1X look like 3x higher memory latency.
The M1 Max also has a crazy number of memory channels, even the older M1 has 4 channels, pro has 8 channels, and max has 16. So you can have many more cache misses in flight, this is part of why the m1 max is 20% faster than the 5950x with twice as many cores.
«The L1 TLB has been doubled from 128 pages to 256 pages, and the L2 TLB goes up from 2048 pages to 3072 pages. On today’s iPhones this is an absolutely overkill change as the page size is 16KB, which means that the L2 TLB covers 48MB which is well beyond the cache capacity of even the A14» [0].
For the default page size on A14 of 16kb, that gives the 4Mb L1 cache coverage (16kb * 256), and 48Mb (16kb * 3072) L2 cache coverage. Off the top of my head only POWER 9/10 (and maybe POWER11) come close with such large TLB's. aarch64 page sizes come as fixed presets, e.g.:
- 4kb, 2Mb, and 1Gb
- 16kb, 32Mb
- 64kb, 512Mb
The OS can choose either of the three, and OS X defaults to the 16kb / 32Mb preset, although I have not seen whether it can simultaneously handle both, 16kb and 32Mb, page sizes.
A garbage collected runtime that has been optimised to make use of large TLB's or large page sizes can exploit the full advantage of the increased TLB depth. Azul JVM comes to mind with their ZGC garbage collector having been heavily optimised for terrabyte scale Java memory workloads.
I am now very curious to see the M1 Max vivisection results that would reveal whether the TLB size in it is even larger (or not).
[0] https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
Here's hoping they put the same in the mac mini. Anyone interested in a linux port join the Marcan patreon, I'm kicking in a few $ a month.
"M1 core which is [...] much wider than both AMD and Intel" The core width you're referring to is the decoder width yeah? As another poster pointed out the M1 has a large reorder buffer as well. Combined with the ability to index a larger L1 cache what I'm getting here is that the M1 can be a lot better about scheduling instructions (and perhaps even running non-interfering instructions in parallel on a given core (is that a thing?)) because the frontend has more power and space to do so.
I guess that efficiency is then a big part of the puzzle on why the increased bandwidth to ram makes such an impact?
Even on the CPU side it can be a big win, in particular on SpecFPRate (a collection of heavy floating point real world codes, not microbenchmarks) Anand has this to say: The fp2017 suite has more workloads that are more memory-bound, and it’s here where the M1 Max is absolutely absurd. The workloads that put the most memory pressure and stress the DRAM the most, such as 503.bwaves, 519.lbm, 549.fotonik3d and 554.roms, have all multiple factors of performance advantages compared to the best Intel and AMD have to offer.
To drive this home compare the Spec2017 FP Rate, the M1 Max gets 81.07, the Ryzen 5950x (high end desktop with twice as many fast cores and a 105 watt TDP) gets 62.27.
So a low power M1 Max with half as many cores and much lower power is 30% faster than AMD's highest end desktop chip. Instead of a desktop size/volume/power, you can get it in a laptop that's 2/3rd of an inch thick.
CAD/CAM, CAE/Engineering, Render farms, Movie making, GIS?