This also keeps me gobsmacked by what a good job Apple has done, designing these bad boys.
This also keeps me gobsmacked by what a good job Apple has done, designing these bad boys.
Yes, exactly. They design chips and market them. They do both things like no other company on this planet.
Most ARM users have much more limited licenses, and can only buy and fab IP blocks designed by someone else (usually ARM).
That is the case of Amazon, the Graviton used A72 cores, and the Graviton II uses Neoverse N1, which are derived from the A76 with more focus on server workloads.
Both are designed by and licensed from ARM.
Apple was one of the original founders of the business entity that licenses the ARM architecture these days. Apple sold the share of that business that they directly controlled.
But the overall result is that specific terms of Apple's license to ARM intellectual property have never been disclosed, and are thought to be quite broad, a license at least on par with architecture and foundry-level licenses.
Not sinister, just an exceptional relationship between Apple and ARM, historically.
An ARM architecture license is not a licensed to use an ARM designed core. It is a license to implement the ARM instruction set on your own non-ARM design. "Architecture" in this case comes from ISA: Instruction Set Architecture. :D
So, Apple, Qualcomm, and Samsung (and Nvidia?) have ARM architecture licenses, and they then go and design their own implementations, from scratch (I don't know if the license lets them use ARM designed bits, you'd have to ask them :D ).
Most companies simply license an ARM designed core and blat that into their SoC - I assume it's cheaper but I don't know if that means the license is cheaper, or the total cost (e.g. including actually designing a cpu) is cheaper.
if only that had been their ingenious plan.
They fixed it in later Touch Bars by adding back a physical escape key next to the Touch Bar. That at least makes it slightly easier to arrange the Touch Bar to not trigger random behaviour :D (apparently I float my little finger above the top left of the keyboard)
Latter for sure, former not so much. Apple is not designing its own CPUs (but ARM does) nor they are physically manufacturing or improving the silicon process (but TSMC does), they build SoCs like many other companies out there. This means that they can't make an advantage over the competitors by:
1. Making the silicon process more advanced
2. Improving CPU microarchitecture
3. Improving CPU instruction pipeline
4. Improving CPU branch prediction or OOO execution
5. Making the TLB page-walker algorithms more efficient
6. Making better cache-coherency protocols
7. Implementing their own bus interconnect which is faster than others
8. Designing better RAM modules
9. Etc. etc. (I could go on and on but I guess the point is visible)
What Apple can do to make a difference is that they can fine tune the SoC microarchitectural details but hardly anything else. By looking at the specs of M1's, the biggest outliers with respect to other chips of the same or similar class currently existing on the market are: 1. RAM and CPUs are physically co-located on the same die
2. RAM is advanced to (LP)DDR5
3. CPU caches are massive (L1, L2 and TLB)
4. There's a separate cluster of CPUs for low-power mode execution
First three things from the list is what is making M1's so fast, and especially the third point. Enlarging CPU caches has been done by the server CPUs like forever and which is the reason why quite old (2014) 2x Xeon system can eat AMD Ryzen 9 5950X for breakfast in some of the multicore workloads I ran (and I mean substantially). Last thing from the list above is what is making this chip more power efficient and this technique is as old as the moment when big.LITTLE was introduced in 2011. In nowadays ARM chips it is known as DynamIQ but the concept is the same: conserving energy by running less power hungry CPUs when workload does not require the full power (which in average user is most of the time).Do you have any evidence supporting this? If you were right, it would seem odd that Apple has consistently delivered significantly better performance, often by multiple product generations, than other ARM licensees since the launch of the A8 in 2014 if ARM was doing all of the real work. Similarly, it seems like they could have saved a ton of money not hiring chip architects if they didn't have anything for them to work on.
Apple IS designing its own CPUs. They have been designing their own CPUs since their A4 processor.
Apple has an Architecture License with ARM [1]. Other companies with the same license include Qualcomm, Samsung, Nvidia, and Broadcom.
These licensees have full architectural freedom with the cores as long as they are ARM ISA compatible. For example, Apple has added custom instructions in their M1 processor [2].
So, an Apple ARM processor core is different from a Samsung ARM core is different from a Qualcomm ARM core.
[1] https://www.anandtech.com/show/7112/the-arm-diaries-part-1-h... [2] https://news.ycombinator.com/item?id=25559145
And apparently in Apple's case, they get to be a little bit incompatible (no nVHE mode, crazy custom ISA extensions, ...)
Which is also obvious proof that they're their own designs, because literally nobody else could or would implement the same Apple-proprietary ISA extensions. Here, we use some of the custom instructions in m1n1:
https://github.com/AsahiLinux/m1n1/blob/main/src/gxf_asm.S#L...
That won't work on any non-Apple core.
That is maybe a little excessive, technically they could probably pay an other architecture license holder to do it for them.
That is not quite true, the A4 is an internal SoC design but commonly "CPU" in an SoC context will designate the cores (the µarch). A4 used a standard and pedestrian Cortex-A8 there.
The A6 is where they started designing the cores in-house (to very impressive and pretty universally praised results, especially for a first iteration), following which they straight nuked the entire field by releasing the first AArch64 core in A7.
My understanding was that the A4 Cortex-A8 was a modified version by Intrinsity which apple bought out, but, you’re right. Apple didn’t start actually designing their own cores until the A6 [1].
[1] https://seekingalpha.com/amp/article/4293287-setting-record-...
What the Apple chips do appear to be is a successful application of the "lead bullets" approach to good engineering. If you compare to Zen 2 it's clear that Apple's designs are not really incomparably better than peers, but they are excellent. With the advantage of better software support and a better manufacturing process than other mobile CPUs they perform really well.
They don't directly compete with CPU vendors but hopefully their new products light the fire under their PC counterparts to build better products. I sure hate feeling trapped in their super buggy software ecosystem.
If you sell chips to device manufacturers for a living, increasing die size cuts directly into how much you make per wafer.
When you sell devices containing your own custom chips, your profit margin does not depend on the cost of fabbing that chip alone.
As a bonus, when everyone else is trying to make a smaller die and cranks up the clocks for performance, you are hugely more power efficient (without giving up performance) when you throw die size at an extra wide instruction decode with a ton of execution units running at a much slower clock speed and have huge reorder buffers so you can track more instructions in flight.
https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
Increasing the cache sizes to be as large as Apple's phone SOC is an extremely recent decision on AMD and Intel's part, as noted at the M1 introduction.
>Last year we had speculated that the A13 had 128KB L1 Instruction cache, Apple has confirmed that it’s actually a massive 192KB instruction cache.
That’s absolutely enormous and is 3x larger than the competing Arm designs, and 6x larger than current x86 designs
https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
> throwing more transistors at the problem doesn't necessarily help power efficiency
Getting your performance from IPC improvements at a lower clock (at the cost of die space) instead of pushing the power envelope to hit ever increasing clock speeds seems to be paying off quite well.
For instance, when power leakage became a huge issue at 20nm, Apple was less affected since they ran their clocks so much lower.
>Apple is something of the exception here, with the 20nm A8 proving to be a solid SoC, thanks in part to their wide CPU design allowing them to achieve good performance without using high clockspeeds that would exacerbate the problem.
https://www.anandtech.com/show/9686/the-apple-iphone-6s-and-...
Keep in mind that most laptops have a 128 bit memory interface as do even high end desktops like i7/i9 and the ryzen 59XX. M1 max and pro are 256, and 512 bits and running at 4266 Mhz!
Performance per core near the best Intel and AMD have to offer, but perf/watt is crazy better.
It's a similar situation to legacy x86 software that doesn't support AVX-512. It still runs, but could run faster if it supported the available hardware.
For example, World of Warcraft is available in an ARM native version that supports Metal and can run at 120 FPS on the 16 inch MacBook M1 Max with the graphics settings cranked up to 10.
https://www.youtube.com/watch?v=JZV3bOpv2Rs
x86 games on OpenGL don't run nearly as well, as expected.
I don't hate Apple because they're a "design and marketing" company; those just happen to be what they're best at. I hate them because they're lazy, they waste money and spend every drop of goodwill they ascertain on locking in consumers, reducing the choice they have inside their ecosystem and raising the wall around their garden. Apple is a company driven by the same ethos that ruined Microsoft 10 years ago, and their reckoning is fast-approaching. The M1 still can't run a ton of software. It is the compatibility underdog, and the x86 developers of the world (at least speaking anecdotally here) simply don't care.
Apple is still not the best chip designer (Nvidia and even AMD nowadays have them beat there), and their marketing is only powerful because they can say whatever they want and people will buy it anyways. You're right on one account though; nobody else markets like they do.
I am not sure what world you live in? Their entry-level chip (the M1) beats workstation-class AMD CPUs like the Ryzen 3700X in most benchmarks, while being passively cooled and using 15W for the whole SOC, while the 3700X uses at least 65W for the CPU alone and requires a large cooling block. All while being passively cooled.
In matrix multiplication workloads, that same entry-level CPU is roughly as fast as a 12 core Ryzen 5900X, while using only 7.5W (the 5900X pretty much goes up to 105W for the same workload).
And all that while delivering stellar battery life.
You can hate Apple for their politics, walled-garden approach, etc. But your comment completely mischaracterizes the engineering feat of M1 CPUs. They have gone from having no presence on laptops/desktops to setting the gold standard for performance per watt.
At any rate, this is all besides the point. Even when Apple was still on 7nm, their iPhone/iPad CPUs were already competitive with Intel parts at much better performance per watt (an also left Qualcomm et al. far behind). It has been clear for a few years already that your point that
to make a CPU that's just barely competitive with the rest of the market
is nonsense.