Where’s the Apple M2?
tbray.org
tbray.org
https://en.wikipedia.org/wiki/Apple_silicon
Second, it's not about clock rate. That is only one small part of the story. It is really about instructions per cycle per core. Apple is killing it on that front and running wider at lower rates is a big part of how they are outperforming in performance per watt while still winning in single core performance. We may see some clock rate increase in an M2, but I suspect their basic design philosophy won't change. It is just working too well.
For the curious, see https://travisdowns.github.io/blog/2019/06/11/speed-limits.h....
That table shows that the M1 has a much bigger reorder buffer, large load and store buffers, huge integer and vector register files, way more branches in flight, etc. By eschewing high clock rates, they are able to really go after massive concurrency at the hardware level in a single core. 7 simultaneous integer operations, 4 simultaneous floating point, multiple load and store. It's a beast.
https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
Of course they will come out with an M2 soon. They've been doing this year after year for over a decade.
How are you counting that?
By the way, yes, both A14 and M1 use Firestorm + Icestorm cores: so I guess that they are the same generation from a technical standpoint.
So was Intel for decades, then they weren't.
Personally I don't see the Intel rut as particularly deep or mucky. Intel has good management, and Gelsinger has deep knowledge of how enterprise customers operate due to his experience. They have a road map for some exciting product releases in the next couple of years, and they dominate their game in terms of market share.
Intel made over $20B profit last year, and semiconductor demand is booming across the board, but they still get trashed by the masses. It's really interesting (if you're into stocks) to compare Intel’s P/E ratio against the rest of the sector. Even the market doesn't think particularly highly of them.
The floating point is going to be an interesting thing to look at - the CPUs made targeting HPC workloads tend to be flops heavy, but the flops tend to be starved for memory bandwidth unless you're doing exactly the best vector processing you can & that fortran can do a great job with.
So you throw in a lot more oomph on the vector side and leave single operation float arithmetic at 2.
Floating point operations of a smaller size of values (more realistically, quarternions or rgba) would be the reason that M1 feels a little bit more snappy when it comes to basic things like text-layout code or graphics images which don't do SIMD very well, but still consume a lot of arithmetic.
I'd suspect that the vertical integration is going to be the secret, because it looks like more profile information of desktop apps going into chip design here.
A similar story is expected of the Graviton series as well, with AWS having a good idea what to build for.
There are 1024 possible Armv8 instructions [1] as opposed to 1,503 x86 instructions [2] and 3,684 x86-64 instructions [3]. There are things x86 and x86-64 can do in a single instruction that would take dozens of instructions to accomplish on Arm.
[1] https://www.csie.ntu.edu.tw/~cyy/courses/assembly/10fall/lec...
[2] https://fgiesen.wordpress.com/2016/08/25/how-many-x86-instru...
[3] https://www.csie.ntu.edu.tw/~cyy/courses/assembly/10fall/lec...
If you have a CPU that breaks one mega x86 instruction in to 100 internal instructions, is that any better than 100 external instructions generated by a compiler?
Theoretically this means that x86 instructions are smaller (with better I$ performance), at the expense of larger/slower instruction decode
It is a lot easier to abstract away complex instructions from x86 (when performance isn't needed) than it is to add physical hardware instructions to ARM (when performance is needed).
Any particular implementation of an architecture is a study in tradeoffs - gates are relatively cheap these days, but not much faster - it makes it easier to throw gates at bigger caches and more cores rather than faster ones - faster clocks likely mean longer pipelines which need lots of well predictable branches to perform which skew an implementation to particular types of benchmarks/workloads
This doesn't really matter, RISC is not a real distinction - it just means "kind of like MIPS". They both have complex address operands and 2-operand vs 3-operand is largely aesthetic.
x86's stronger memory model is what matters most since it can reorder memory accesses less often.
AMD has admitted to being at our very near the upper limit for x86 decoding. With arm Apple can have and reliably feed more decoders from the incoming stream of instructions. They have built their whole chip around being able to extract parallelism from the huge instruction window.
That cost is that it's larger. At the same power AMD can stuff twice the cores, with similar though lower single core performance. That also means lower clock speeds, and there are less instruction level guarantees you can rely on thus somewhat more complexity is necessary.
And no, it's not really a beast. It's competitive.
Note that the “TDP” is meaningless for actual performance/watt comparisons with a controlled parameter due to the variation in this term. It’s just a marketing term.
The Ryzens on 7NM consume much more power than Apple’s Big cores did on the A12 and A13, both of which were on TSMC 7NM. They are also both competitive with the single-core scores of the Ryzen Zen 2 or 3 cores, if not equivalent while being older architectures than AMD’s.
The M1’s Big [Firestorm] cores also consume less power and achieve more performance than the Zen 3 core.
It’s safe to say Apple’s architecture is largely superior, outright.
I'm sorry, but you're comparing power draw of a workstation "X" AMD chip with a laptop chip. It's simply not a honest comparison. You must compare the efficiency of mobile chip with a mobile chip. When you do that you find similar efficiency.
I don't understand why no one is posting actual apples to apples comparison. Every time a comparison is posted its either comparing to a workstation chip to find power efficiency even if they're tuned to be power inefficient, or comparing to Intel CPUs only, etc...
Laptop processors from AMD use less than half the power per core of workstation processors while sacrificing only a very small performance gain, due to the use of a different lithography for I/O.
Apple's architecture is simply not superior. If it was, we wouldn't be making these contrived comparisons, and we wouldn't even be comparing 5nm chips to 7nm chips.
The IO die only adds like ~15W btw, yes, I'm aware it's on a GlobalFoundries node.
Lastly, AMD's mobile chips throttle down to fairly low clocks and in AMD's case, low to modest performance when not plugged in, it's why you keep hearing tale of the great battery life on Zen 3 laptops.
That leaves out the most important thing that enables all of that, the 8-wide symmetric decoder that can feed those. x86_64 cpus only have 5-wide ones, and only the first of those can decode the multi-uops instructions, and even worse there are even more complex instructions that are microcoded.
They told us they were on a two-year transition, and we have until WWDC next year at the earliest to finish that transition. And looking at what happened during the first year, it's pretty clear they meant that each machine would be updated once during the transition.
The Macbook Air, Mac Mini and 13-inch Macbook Pro were obviously the machines deemed to skip a redesign, and were released first. The 24-inch iMac was redesigned, and I suspect every other computer will get a redesign to go with their new chip.
I suspect this Fall, after the new iPhones unveil a new set of cores, we will see those cores used to build a new chip that goes in the machines still due for an update: the big iMac and the big Macbook Pro. The M1 machines will not be updated, and we will only see a Mac Pro at WWDC 2022. Then who knows?
It's in Apple's best interest to stop selling x64 arch computers as fast as possible.
The logical order would be to do a blazing fast low end chip first in the Air/Mini (tbh I'm surprised they diluted the mbp brand with an iPad chip), with a higher end MBP/iMac chip less than a year later. The iMac Pro was already discontinued and the Mac Pro is vastly less important wrt the architecture transition and can easily be done last.
I expect M>1 iMac and/or MBP in September.
If Apple is dealing with any surprises relating to the silicon, those would most likely be security issues discovered/reported in the M1 platform after launch. Only Apple knows what's on that list – it's definitely not empty – and whether any of them must be fixed before M1X devices ship.
The history here might not be representative because many of those delays were caused by Motorola/IBM and Intel struggling to deliver chips in the expected volume or thermal budget. Since that was one of the motivations for Apple to make their own designs it should be less pronounced going forward, especially without Johnny Ives pushing the limits so hard.
Get a load of this guy. Can't even wait even a full year, he wants his CPU revolutions every 4 to 8 months.
If the big Macbook Pro is only released this Fall or next Spring, it will be business as usual for Apple.
Same goes for an update to the M1 Macbook Air. I expect that machine to be updated next year at the earliest, with something akin to an M3.
And they are banking hard on unified memory model in all sorts of marketing zbut it's unclear if it's the cause or effect of ring unable to use external GPU.
The M1 was hailed as an entry level chip, and merely the start of greater things to come. If that were true, logic would hold that they would have a higher end version of the existing M1 chip available by now.
It's not a different model architecture.
A new MacBook Pro with M1X will come out this fall according to credible sources.
Hailed by whom? I doubt you'll find a single instance of Apple calling M1 an "entry level chip".
It looks like this is the blogosphere getting too high on their own supply, huh?
Apple's chip design team is working on M3 if not M4 right now.
The failed release of shaped batteries in the 2016 MBP [0], and the jet-falls-off-the-aircraft-carrier release of AirPower. [1]
I generally agree Apple makes plans in advance and follows through on them. However, the company seems to push very hard on some product release deadlines and sometimes they miss.
[0] https://www.bloomberg.com/news/articles/2016-12-20/how-apple...
[1] https://www.bloombergquint.com/technology/apple-cancels-anti...
Image import is generally I/O bound, so not a good fit for CPU comparisons.
The GPU issue is relevant, we'll see how the M2 does there. Will Apple need a discrete GPU to compete?
All that said, the M1 is most impressive in terms of performance/Watt. We'll see how the M2 holds up against Threadripper/Epyc once the Mac Pro refresh is done and it's benchmarked with many CPU/GPU bound pro workloads.
The M2 should be nice for Macbook Pros though. Looking forward to it!
I can already hear the talking heads now
"Most users don't need that power(but apparently they need the power of the m1?)"
If I know Apple, it will be medium end at extraordinary prices. Post purchase Rationalization will cause users to praise it regardless.
There isn't evidence that Apple's team has run out of architectural improvements however, so I do think performance gains are still out there. Plus there's always the possibility of going to smaller semiconductor processes.
Not that I would have expected the step function they did manage to pull off either…
It's not a very interesting article but it specifically talks about the relative non-importance of that particular benchmark, beside the results being largely a wash:
I sorely miss the benchmark I saw in some other publication but can’t find now, where they measured the interactive performance when you load up a series of photos on-screen. These import & export measurements are useful, but frankly when I do that kind of thing I go read email or get a coffee while it’s happening, so it doesn’t really hold me up as such.
To date, I haven’t heard anyone saying Lightroom is significantly snappier on an M1 than on a recent Intel MBP. I’d be happy to be corrected.
Plus, why the stupid monitor limitation? A "Pro" MacBook needs to be able to have 2-3 monitors with no compromises.
I'm currently using 40GB and I rebooted on Monday.
As it is, if you're running Docker Desktop for Mac, you're spinning up an entirely separate (internal) VM with macOS filesystem syncing for each Docker container you run.
All I want is a 14" laptop with an M1 and 32GB of RAM. And the extra RAM is just so I can comfortably allocate RAM to Safari tabs.
This is daft. Chips are not something that you can finalise the design of and have in customers hands tomorrow, next week, or next month.
Whatever pro-level AS device Apple ends up shipping, the design of it was finalised by the end of last year at the latest. What's far more likely is they either never planned to ship it ad WWDC, or (more likely) they're affected by the same market dynamics as everyone else in terms of the ongoing semiconductor shortages.
The idea that Apple made the m1, and somehow got stuck now is a bit silly. Most likely, they know what they are doing. They might hit unexpected problems 3-4 years from now, but not right now.
The m1x and m2 will only be incremental steps forward. That’s how the A-series have progressed, and that’s fine.
It’s possible the arguments in the rest of the article are correct, but this isn’t. Hardware cycles take a while and the design for the M1X or M2 was surely settled upon some time ago, probably even before the official M1 release.
A whole eight months?
How often does he want Apple to release new products? What's his problem?
I disagree and I think that based on what we've seen with the current constraints (fanless etc.) actually Apple can come out with some very competitive products in the pro/desktop/larger-laptops space.
In fact, I also think that probably the GPU will be the main differentiator between low-end and high-end Apple silicon models.
1. Apple could take the M1 design and just increased the number of Firestorm CPU cores (along with everything else). It would be an M1X, along exactly the same lines as the A10X and A12X.
2. Apple could create a next generation design with the eight or more of same Firestorm-next cores we will see in the A15 Bionic. Call it the M2.
The M1X would have been easy enough for apple, and I think if it existed, we would have seen it by now. But it would have suffered from many of the same flaws that the M1 suffers from. The inability to drive more than two displays being a notable one.
I think apple decided to skip the M1X and go straight for the M2.
But it's a little too early for the M2 to launch. As it will be using the same cpu and gpu cores as the A15, it kind of needs to launch at around the same time at the earliest; Maybe a month or two earlier as they don't need to stockpile as much silicon for the M2s.
My prediction is that we might see an M2 announced in August or September.
No? Larger chips generally also have more I/O.
But try to drive two 1080p or 720p monitors. Impossible.
The issue is the M1 only has two CRTCs to drive the video timings.
Unsure what the point is here.
And there are plenty of workloads which are easily parallelizable across cores. Like compiling things. Or just running multiple programs at once.
Which you can see by typing “ps”.
I don't see why people complain its only one benchmark. to import 100MB of image data invites all of the IO, Cache and CPU to play along because this is encoded data: it has to be mapped out of one form, transformed into another, as a stream of data and its a lot bigger than either a single fetch from memory, or a single bus transaction. Its ameneble to parallelism within some limits depending on the nature of the encoding. It's also a real-world test.
This seems to be throwing the baby out with the bathwater. People have mostly been saying that everything else is snappier on an M1, app launches, task switching, etc.
Lightroom keeps claiming this on new releases, and have done for years, and users keep finding very little improvement. Adobe clearly wants users moving to the non-Classic cloud.
> Anyhow, it’d be really surprising if Apple managed to get ahead of GPU makers like NVidia. Now, at this point, based on the M1 we should expect surprises from Apple. But I’m not even sure that’d be their best silicon bet.
I would be very surprised, competition is intense and there is a lot of money on the table for anyone with a better architecture. Nvidia would compete for any acquisitions. Ultimately does Apple even need to compete with Nvidia? They don't care about gaming, just desktop graphics and local ML inference.
1) The best silicon designers in the world have hit a wall and can’t improve on a processor they shipped to production a year ago.
2) There is a massive, global-scale supply chain disaster that is largely hidden from the end consumer which is secretly driving every major manufacturer insane.
It seems that a bunch of industry “insiders” want to go with option 1; personally, I’ll take the second option.
All I’ve heard is that the mini LED screens for the next (M1X) MacBooks haven’t been available in volume/to spec quite yet.
I’m really not thinking there’s any real delay of concern here yet. Not enough to write the above article and speculate somewhat needlessly.
This is only true for editing still images. If you're editing video (with tools like Davinci Resolve - which is free and excellent), 16GB is going to be quite inadequate.
And I've seen reports that the M1 can handle RED RAW 4K footage with just 8G of RAM, and only stutters on 8K. Snazzy Labs even claims that it actually scrubs the timeline smoother than this $10,000 Mac Pro ever did, the only failing is 50% longer render times: https://www.youtube.com/watch?v=eY-S9EuJ5Xs&t=474s
So this once again proves that the M1 defies our standard expectations and measurements of computer processing power.
Sheesh, 8 months is not a long time as far as these sorts of things go. Wasn't the M2 not expected until next year at the earliest? Given the current semiconductor production issues that might even be optimistic.
While I don't buy this at all, I don't even want a chip that is perceptibly faster than the M1. Instead give me other "pro" features like more than 2 USB ports, 32GB RAM, multiple external displays, external GPU support, boot camp, larger SSD capacity.
m1 is a chip designed to fit into their current lineup. it makes sense to launch m2 with redesigned hardware or after a hardware redesign. ergo, the M2 will need to fit in the chassis of the new iMac.
hardware redesign + new in-house chip = a lot of time
I'd be amazed if this even drops this year instead of early 2022 with how long apple's release cycles are. think they'll prob roll out the smaller MBP, the Air, the mini redesigns in the fall/winter, while keeping the current 16" lineup. they'll launch the 16" and larger iMac at next summer's WWDC.
1) cross execution caching of unchanged segments of code: if the ram state doesn't change across runs for a set of instructions, could execution be sped up by caching that ram state with os and silicon support for applying it?
2) assured computing: cryptographic guarantees about exactly which instructions executed, on which physical machine with what inputs
3) hardware support for accelerated vms emulating other architectures
4) more efficiency gains by having more stuff like the low power cores for light workloads
5) more of the cool specialty cores for accelerating ml/dsp/linalg
6) better in-silicon multitenant separation for security
7) in-silicon support for reducing memory costs of running multiple versions that are mostly similar of libraries (hardware support for containers)
are just a few... apple is in a great place because their vertical integration makes some of these possible. (and the good stuff of course would eventually make it to commodity hardware)
LOL. Has the author even spoken with actual photographers? Also, Lightroom Classic? Seriously? Lightroom CC is blazing fast on an M1. I know because it has been my go to real-life benchmark with all the M1 Macs I tested so far (all of them). You should of course compare them with similarly priced Intel Macs. But anyway, try working on 1000+ Canon Eos R5 raw files on a MacBook Pro M1 and the latest MacBook Pro Intel, then tell me which one’s faster. I’m not sure I would be able to hear you over the noise of the Intel Mac’s fans, though.
This “perception of speed” theory doesn’t hold up to explain the M1X/M2 delay. I definitely believe TSMC’s bottlenecks are a much more probable explanation.
All they need to do is support external RAM so they can make a machine with up to 128 or 256gb of memory. It's easy enough, but for their reasons they have just chosen not to do it yet.
The best analogy I can think of is to imagine if epoll or dbus were hardware-accelerated, and every year a silicon update and a kernel update added more or improved acceleration for more and more components.
We don’t yet know what the speed gains possible from reducing OS overhead are, but I bet Apple does, and I bet it doesn’t require frequency bumps at all. If someone has compiled vmlinuz into FPGA tapeout somehow, that would be a good point of comparison. (“Inconceivable!”, except not so much nowadays..)
[I can’t find the tweet linked in the past few weeks about CPU-accelerated ObjC calls, but there’s an HN discussion about it that’s worth reading.]