Competition is great.
Competition is great.
You can already get a usable desktop on an M1 Mac, if you really want to. The CPU is fast enough that you don't absolutely need graphics acceleration, and Wifi, USB, and display output all work.
The big missing piece is the GPU, but Alyssa has a fully-custom user-space implementation for macOS that largely works.
All of this is SDRAM, the S stands for synchronous and means that the memory is driven according to fixed timings in relation to a fixed bus clock. All LPDDR4X-4266 parts with the same timings perform exactly the same, whether they are soldered to the interposer or are 5 cm away on the board.
"M1 core which is [...] much wider than both AMD and Intel" The core width you're referring to is the decoder width yeah? As another poster pointed out the M1 has a large reorder buffer as well. Combined with the ability to index a larger L1 cache what I'm getting here is that the M1 can be a lot better about scheduling instructions (and perhaps even running non-interfering instructions in parallel on a given core (is that a thing?)) because the frontend has more power and space to do so.
I guess that efficiency is then a big part of the puzzle on why the increased bandwidth to ram makes such an impact?
Even on the CPU side it can be a big win, in particular on SpecFPRate (a collection of heavy floating point real world codes, not microbenchmarks) Anand has this to say: The fp2017 suite has more workloads that are more memory-bound, and it’s here where the M1 Max is absolutely absurd. The workloads that put the most memory pressure and stress the DRAM the most, such as 503.bwaves, 519.lbm, 549.fotonik3d and 554.roms, have all multiple factors of performance advantages compared to the best Intel and AMD have to offer.
To drive this home compare the Spec2017 FP Rate, the M1 Max gets 81.07, the Ryzen 5950x (high end desktop with twice as many fast cores and a 105 watt TDP) gets 62.27.
So a low power M1 Max with half as many cores and much lower power is 30% faster than AMD's highest end desktop chip. Instead of a desktop size/volume/power, you can get it in a laptop that's 2/3rd of an inch thick.
I'll go first.
Personally when talking about memory latency, I want to measure time to main memory, without TLB thrashing. Not that TLB latency isn't interesting, but it should be separately quantified.
I've written a few microbenchmarks in this area, but I get very similar numbers to the Anandtech R per RV prange. Which puts the latency to main memory at around 35ns looking like it might flatten out at 40ns. Yes, full TLB thrashing is up at 110ns or so, but that's not a usual use case. If it is, at least under linux, you can switch to 1GB pages if that's important to you.
R per RV prange numbers on the Ryzen 9 5950x is 65ns or so.
So sure the TLB worst case is higher latency on the M1, but then you have to figure out how big the TLB is, and how often that's an issue if you want to know the real world performance impact.
The i9-12900K with DDR5 gets 30ns on the R per RV prange, and 92ns on full random (tlb thrashing).
Even assuming the worst case TLB behavior the m1 max is 111ns and the 5950x is 79ns or 1.4x higher. On the Intel side is 111ns vs 92n, 1.2x higher.
Finding it hard to find any numbers that make the M1X look like 3x higher memory latency.
The M1 Max also has a crazy number of memory channels, even the older M1 has 4 channels, pro has 8 channels, and max has 16. So you can have many more cache misses in flight, this is part of why the m1 max is 20% faster than the 5950x with twice as many cores.
Here's hoping they put the same in the mac mini. Anyone interested in a linux port join the Marcan patreon, I'm kicking in a few $ a month.
«The L1 TLB has been doubled from 128 pages to 256 pages, and the L2 TLB goes up from 2048 pages to 3072 pages. On today’s iPhones this is an absolutely overkill change as the page size is 16KB, which means that the L2 TLB covers 48MB which is well beyond the cache capacity of even the A14» [0].
For the default page size on A14 of 16kb, that gives the 4Mb L1 cache coverage (16kb * 256), and 48Mb (16kb * 3072) L2 cache coverage. Off the top of my head only POWER 9/10 (and maybe POWER11) come close with such large TLB's. aarch64 page sizes come as fixed presets, e.g.:
- 4kb, 2Mb, and 1Gb
- 16kb, 32Mb
- 64kb, 512Mb
The OS can choose either of the three, and OS X defaults to the 16kb / 32Mb preset, although I have not seen whether it can simultaneously handle both, 16kb and 32Mb, page sizes.
A garbage collected runtime that has been optimised to make use of large TLB's or large page sizes can exploit the full advantage of the increased TLB depth. Azul JVM comes to mind with their ZGC garbage collector having been heavily optimised for terrabyte scale Java memory workloads.
I am now very curious to see the M1 Max vivisection results that would reveal whether the TLB size in it is even larger (or not).
[0] https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
"has a big reorder buffer", I'm interpreting this as "one of the notable advantages of the M1 is it's ability to be more clever about processing instructions out of order to maximize resource utilization". Is that about right?
If you’re really interested in this it might be worth finding a copy of the Patterson and Hennessy book. It’s a big read and expensive (but older versions are on the internet archive [1]) and covers all these design issues in quite a lot of detail.
[1] https://archive.org/details/ComputerArchitectureAQuantitativ...
CAD/CAM, CAE/Engineering, Render farms, Movie making, GIS?
Suddenly Intel and AMD have to keep an eye not just on each other, but also on the various ARM designs targeting microcontrollers up to supercomputers.
For reference, in Zen 2/3 a chiplet (of eight cores) is limited to around 25 GB/s write, 50 GB/s read.
Edit: The anandtech test has memory B/W numbers and gives 100 GB/s for a single core and around 220 GB/s for all cores, which is extremely high, but also not the full memory bandwidth.
Marketing brags about memory bandwidth based on clockspeed * bus width = 400GB/sec. Getting 60% of peak on some memory bandwidth benchmark is pretty common. Try McCalpin's stream benchmark on your platform of choice to verify. I suspect you'll find similar on Intel or AMD on desktops, laptops, or servers.
Not sure I buy the memory bound argument, sure peak performance will not change, but worst case performance can be caused by cache misses. So I expect that the new MBPs will be more evenly fast than the x86-64 competition, even while doing stressful things. I don't currently need a $3k laptop, but am hoping they ship a mini with a M1 Max or Pro.
To quote anandtech: The fp2017 suite has more workloads that are more memory-bound, and it’s here where the M1 Max is absolutely absurd. The workloads that put the most memory pressure and stress the DRAM the most, such as 503.bwaves, 519.lbm, 549.fotonik3d and 554.roms, have all multiple factors of performance advantages compared to the best Intel and AMD have to offer.
So an apple M1 Max is 19% faster on SpecFP (floating point applications) as a Ryzen 5950x which has twice as many fast cores (16 vs 8) and runs at 105 watts TDP. That's pretty amazing in my book.
What other kind ofARM hardware can we buy beside Apple laptops?
Thinkpads are notoriously well made. 'shitty plastic' is ultra durable. Macs are great if you are ok being locked into the Apple ecosystem
In real world use I don't notice any difference in restrictions from my Ubuntu server.
You can also buy RISCV if you want something more open and not ARM.
Or phones or tablets.
Edit: first thing I read was wrong, turns out that these are Mac Pro level processors, see below comments for more details. Thanks everyone for the clarification.
Jade C "Chop" is the Pro, Jade C Die is the Max.
jade 2c is two Maxes. Jade 4c is 4 maxes.
https://twitter.com/siracusa/status/1395706013286809600?lang...
But they are expected be even higher scaled iterations of Apple Silicon. In other words, while M1 had up to 4 performance (P) cores, 4 efficiency (E) cores and 8 GPU (G) cores, the M1 Pro and Max scaled up to as high as 8P, 2E and 32G.
The Jade 2/4c designation is rumored to go even higher on P and G cores, which means it's much more likely to end up in a Mac Pro than the Macbook Air.