Graph conveniently omits AMD who now beat Intel using the same ISA but different process.
Graph conveniently omits AMD who now beat Intel using the same ISA but different process.
M1 apparently has much lower power consumption.
So something else is going on other than “Apple is willing to spend more to make cores wider”.
Even compared to similar manufacturing processes ?
To me the innovation is using phone-type SoC technology.
A lot of the power (and die area) demand in modern chips is in the interface circuitry to drive the few centimetres of PCB trace between ICs.
Put the RAM in the same package as the CPU and you eliminate that interface. That gives you die space and power budget for big buffers and lots of parallel execution.
Interface transistors (and wires) are many times larger than the core compute logic transistors (and wires), so there is much more than a one-for-one gain.
"ROB" and "OOO" might have gotten mixed together here.
OOO = Out-of-order. Refers to the fact that the M1 can decode instructions in parallel.
ROB = Re-Order Buffer. Refers to the stage where the parallel instructions get put back "in-order" and "retired."
[1] https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
The other factor is Apple is using a "better" process node (TSMC 5nm). I put it in quotes because Intel's 10nm and upcoming nodes may be competitive, but Intel's 14nm is what Apple gets to compete against today, right now.
Intel has been defeated in detail.
Intel's 10nm node is out, I'm typing this on one right now. It's competitive in single-core performance against what we've seen from the M1. Graphics and multi-core it gets beat though...
Or do you mean what Apple used to use? (edit: the following is incorrect) It's true Apple never used an Intel 10nm part.
EDIT: I was wrong! Apple has used an Intel 10nm part. Thanks for the correction!
There are still no 10nm parts for the desktop or the high-end/high-TDP laptops anyway afaik.
Basically computers that start looking and acting more like clusters. Smarter memory and caching, zero-copy/fast-copy-on-mutate IPC as first-class citizens. More system primitives like io_uring to facilitate i/o.
The popularity of modern languages that make concurrency easy means more leveraging of all those cores.
That's only for 32-bit ARM; for 64-bit ARM, the instruction size is always constant (there's no Thumb/Thumb2/ThumbEE/etc). It won't surprise me at all if Apple's new ARM processor is 64-bit only (being 64-bit only is allowed for ARM, unlike x86), which means that not only the decoder does not have to worry about the 2-byte instructions from Thumb, but also the decoder does not have to worry about the older 32-bit ARM instructions (including the quirky ones like LDM/STM).
That would also explains why Apple can have wide decoders while their competitors can't: these competitors want to keep compatibility with 32-bit ARM, while Apple doesn't care.
With more instructions decoded per clock and a larger reorder buffer, the core should be able to keep its units busier, i.e. not wasting energy without producing valuable output.
This efficiency gain of course needs to outweigh the consumption of the additional decoders. This part is easier with ARM as decoding x86 is complicated.
In addition to higher unit utilization, the increased parallelism should also be an advantage in the "race to sleep" power management strategy.
Apparently ARM supports 4KB, 64KB and 1MB (I just googled that)
Also maybe all the x86 crud might be finally taking its toll (and not just the ISA but all the cruft that got on top of it for 40yrs+ like legacy buses and functionalities, ACPI, or even the overengineered mess that's UEFI)
EFI doesn’t seem that bad either. Maybe something like uboot would suit you better? But the reason we have those nice graphical bios utilities now is because of EFI.
Hence why laptops and 2-1 hybrids are basically back to the days of vertical integration of 16 bit computers, whose the PC was the only exception.
Not really, PC laptops and hybrids are pretty standardized. A slim case doesn't really make it vertical integration.
It is getting more and more integrated though which makes it hard to replace parts. But that has more to do with greed and size requirements than vertical integration, a soldering iron can amend some of that. Driver situation isn't by any means something new or necessarily an indication of vertical integration either.
As for the rest, nothing that regular consumers would ever bother with.
We are talking about vertical integration are we not? Don't see how any of that is relevant then.
And since you did your research the wider public would appreciate to learn in what consumer shop one can get such out of the box experience.
And it's a poor analogy because driver licensing is almost a straight consequence of how dangerous cars are, but the openness of an SoC does not directly follow from how fast it is.
The DMV is a revenue source for most states. Same with traffic cops. They both only have a tenuous link to traffic safety.
1. Power efficiency (read: battery life) 2. Performance 3. Openness
99% of consumers only care about 2 of those. The more that time goes on, the more it becomes clear that prioritizing openness, in particular for hardware, is a direct trade-off against building a great consumer product.
They are able to shovel transistors at the core because there are only 4 of them. AMD on the other hand is packing 64 cores into a package, so each core has to make do with less.
AMD in particular wasted an entire generation of Zen by having too few BTB entries. Zen1 performance on real software (i.e. not SPEC) was tragic, similar to Barcelona on a per-clock comparison. They learned this lesson and dramatically increased BTB entries on Zen2 and again for Zen3. But the question in my mind is why would you save a few pennies per million units by shaving off half the BTB entries? Doesn't make sense. They must have been guided by the wrong benchmark software.
What 'real software' are you thinking of? Anything in particular? Just curious, not looking to argue.
(sorry I changed my comment around the same time you replied)
So they are working on a smaller process node than their competitors, and they used their market position to ensure that they will be the only ones using it for a little while. That gives them a significant leg up.
Skylake was the "tick", then they couldn't "Tock" to 10nm, so they made 14nm+, then 14nm++, then 14nm+++... Intel's been stuck on the same process for half a decade now.
Even worse - the same architecture as well, which is why they have to backport their 10nm Ice Lake architecture to 14nm to even get anywhere.
That could not have come at a worse time for a company like intel.
[1] https://www.reuters.com/article/us-intel-ceo-idUSKBN1JH1VW
Another part of it is just that designing a CPU with good single threaded performance is hard and costs a lot of die area and no one else in the ARM world was incentivized to do this - who needs a smartphone with that kind of performance (I know, I know - all of us, now)? But Apple also makes IPads which are kind of a bridge device - mobile and battery-powered, but also big and something people might want to do real work on.
I’m also going to say that I think Apple is still alone among the phone manufacturers in recognizing that smartphones are software platforms, not simply hardware devices. I think they have a special appreciation for the enabling power of a more capable CPU.
The 5 Watt iPhone CPU almost beats the 45 Watt AMD CPU. And the M1 will then simply obliterate the x86 incumbents.
Why? The real reason is that Apple had/has the real advantage. It's size, resources, money, strategy, know-how.
It's bigger than Intel. It has a better corporate culture than Intel, it has a vertically integrated market completely safe from any outside disruption for years, and will be for years. It had the opportunity to do this switch, so it had the motivation to develop this capability.
Apple was able to hire the best semi designers plus benefit from the ARM platform.
AMD is still too small and relatively spread thin - trying to cover multiple segments of the general market - compared to Apple, which basically has this one chip in two versions (mobile and ultra mobile).
Apple has and had the iPhone. Apple started using their own chips, the A5, about 10 years ago. (In 2011 March, starting with the iPad 2. And they used ARM before that too.) And since then Apple optimized the "supply chain", incremental steps, but when you are Apple every step can be almost revolutionary (and year after year Steve told exactly that to the believers, and it's basically true, just not on actual feature front, but on hardware and systems level - eg. Gorilla glass for the phones, power efficiency, the oversized battery for the MBP, milling, machining, marketing, design, and so on).
Of those two advantages, the process one is a lot more significant. Their chip loses in single-threaded performance to AMD's newest, which are still one node behind Apple, but appears to be more power-efficient. I fully expect that once AMD gets on the matching node they will take the crown back, even in a laptop power envelope.
Everyone else with the immediate capability are professional ball-droppers.
Apple was actually the one designing those cores in the Exynos - up until and including the Exynos 4000. That was when Apple decided to sever ties with Samsung and stopped designing cores for them and Samsung had to go back to stock ARM designs with the Exynos 5000. That was why the 5000 series sucked so badly - Samsung lost access to those Apple/Intrinsity juiced cores. It’s also what prompted Samsung to open SARC in Austin.
Motivation to do so.
Intel had none.