Wow.
Wow.
Does anyone know of more technical reasons for Apple's ARM processors outperforming everyone else's by such a large margin, and for such a long period of time? Seems like there's some fundamental difference in what Apple is doing, and I'd love to read more about it.
People still seem to like writing their own cool malloc, but memcpy not so much.
Things like malloc are quite a bit more complicated, and more workload dependent so there's still some opportunity specializing an implementation there.
I also thought kernel devs would be working against processor ISAs, not hardware-specific details beyond the ISA.
EDIT: But I think another part of it is just being willing to throw more transistors at the problem than Android phone SoC manufacturers are and also that their higher income makes spending more on engineering make sense.
As for Minecraft Pocket Edition, isn't that written in native code anyways? So I'd expect it to perform better on Apple hardware, assuming the hardware is actually the bottleneck for performance.
Mostly, A-series chips have enormous L1 cache, great cache hierarchy and management, and very low memory latency. A12 specifically seems to have included an almost total redesign of the cache hierarchy.
I'm sure there are many more reasons their designs significantly outperform competitors, but I've not seen any more publicly available analysis.
Has this always been the case compared to contemporary Qualcomm/Exynos/etc. SoCs? Not implying your statement is wrong here; all I know is that the A-series chips have had a big performance advantage for a while now and I haven't read more detailed analyses in the past that may have given hints as to why.
Also, how difficult is cache hierarchy/management to get right? For something as fundamental to good performance these days as cache, I would have expected the major players to be on more or less the same playing field.
Apple optimises chip performance
Others optimize chip size
Of course, this iteration apple have beaten almost everyone else to 7nm, so the difference is much more dramatic.
Yes, PA micro was the place where the "last of Mahicans" of US chip industry were.
>Also, apple have pursued 2 (very) fast cores whereas other chips often have 4.
Yes, because people into app development are as web developers, and the word "mutex" gives most of them a panic attack.
Android style java should've been more multi-threading friendly, but that does nothing about people not utilising them.
> this iteration apple have beaten almost everyone else to 7nm, so the difference is much more dramatic.
Yes. I remember how Mediatek beaten Apple to 10nm thanks to them being a Taiwanese company, but nevertheless "ruined it all" with their helio x30's design being designed with more marketing considerations than engineering ones. Their marketing guys couldn't wait to announce "hey we have 2 more cores than you Qualcomm!"
Yes, that's my understanding. I've seen this discussion last year and the year before. And every time the answer seems to be caches. Cache memory is expensive. Apple seems willing to pay more for the SoC in order to have an overall experience that lets them get away with the high prices.
From what I read, Qualcomm would not be able to sell at volume an equivalently performant SoC.
Android manufacturers have more competition, and have to address the low end of the market too. Their money, and attention, is spread in a wider swath.
Samsung comes to mind, surely it's big enough to produce high-performance phones that can compete architecturally with Apple's SoC as well as addressing the developing nation / low-cost phone market?
Sure, yes, but not at a volume that allows them to create a processor that competes with iPhone only on their flagship model. Also, they can't charge $999 for their flagship. The average sales price for an iPhone is higher than the flagship model at Samsung.
https://www.gsmarena.com/analysts_average_selling_price_of_a...
From everything I can tell and looking at the numbers, the high end Android market is minuscule.
The tighter focus, combined with the fact that Apple rakes in more cash to spend on R&D, is why their engineering team is able to win out here.
>Monsoon (A11) and Vortex (A12) are extremely wide machines – with 6 integer execution pipelines among which two are complex units, two load units and store units, two branch ports, and three FP/vector pipelines this gives an estimated 13 execution ports, far wider than Arm’s upcoming Cortex A76 and also wider than Samsung’s M3. In fact, assuming we're not looking at an atypical shared port situation, Apple’s microarchitecture seems to far surpass anything else in terms of width, including desktop CPUs.
https://www.anandtech.com/show/13392/the-iphone-xs-xs-max-re...
Apple first moved to wide CPU designs with the Cyclone CPU core found in the A7 SOC first used in the iPhone 5s.
>With Cyclone Apple is in a completely different league. As far as I can tell, peak issue width of Cyclone is 6 instructions. That’s at least 2x the width of Swift and Krait, and at best more than 3x the width depending on instruction mix. Limitations on co-issuing FP and integer math have also been lifted as you can run up to four integer adds and two FP adds in parallel. You can also perform up to two loads or stores per clock.
On a Xeon v3 core, SPECint averages below 2 instructions per cycle: https://www.researchgate.net/publication/322745869_A_Workloa.... How does Apple beat Intel on branch integer benchmarks like 403.gcc by a factor of two per clock?
Mostly I'm curious about how complete the bypass network is on their functional units and if execution is clustered like the POWER8. The width doubling in the A series does remind me of the POWER 7 to 8 transition.
Renaming is also apparently a major constraint on design width in many cases but I'm not so familiar with that.
What does "FO4" stand for here? Googling it yields "Fallout 4", which definitely isn't right, and I'm not sure what other keywords to tack on to get the right result.
> Fan-out of 4 is a process-independent delay metric in CMOS tech.
The wrap-up slide at the very end has a bit of a rundown.
http://people.duke.edu/~bcl15/teachdir/ece590_fall14/Present...
You can check that part at least. Isn't it LLVM?
I think it's unlikely they'll update their flagship Mac to a new CPU architecture without notifying developers first. On day one, all existing apps would perform poorly, and that's not how you promote a new top-of-the-line system.
”This also gives us a great piece of context for Samsung’s M3 core, which was released this year […] Here the Exynos 9810 uses twice the energy over last year’s A11 – at a 55% performance deficit.”
Doesn’t that mean the A11 already is three times as efficient as recent Android cores? (The Samsung M3 was in January’s Hot Chips)