Arm Announces New Mobile Armv9 CPU Microarchitectures
anandtech.com
anandtech.com
The Cortex A55 (announced 2017) was also a very minor upgrade over the Cortex A53 (announced 2012). These small cores are truly the baseline performance of application processors in global society, powering cheap smartphones, $30 TV boxes, most IoT applications doing non-trivial general computation and educational systems like the Raspberry Pi 3. If ARM’s numbers are near the truth (real-world implementations do like to disappoint us in practice), we’re going to see a nice security and performance uplift there. Things have been stale there a while.
If anything this announcement means that in a few years time, the vast majority of application processors sold will have robust mitigations against many important classes of memory corruption exploits [1].
[1]: https://googleprojectzero.blogspot.com/2019/02/examining-poi...
For what it's worth I'm still daily driving LG G2 and it's legit still fast, sure sometimes it's a bit sleepy but it's still more then usable!
Marques Brownlee aka MKBHD and even he admitted that's not the case anymore.
Good phones are getting more and more expensive. A flagship used to cost $700 5 years ago, now it's $1000+.
Cheap phones are getting good, but not at the rate that good phones are getting more expensive.
> For what it's worth I'm still daily driving LG G2 and it's legit still fast, sure sometimes it's a bit sleepy but it's still more then usable!
A quick Googling tells me your phone is from 2013 and can only be updated to Android 5. Your phone is either:
* horrendously out of date and insecure
* using a custom ROM (which I, for one, and many others probably too, don't want to do)
* not used for any kind of internet facing activity (another thing that I and others don't want to do)
Also, you probably have a very low bar for usability and speed. With modern OSes and sites, that phone would probably be extremely slow and laggy for my (and many others') usage.
Your phone uses a Snapdragon 800 which probably has 10-15% (maybe 20-25% if I'm being generous) of the performance of a modern Snapdragon. Phones are not desktops/laptops, where performance levels plateaued in 2005.
I wonder how much the inflation affected the pricing of the phones.
Also you can get Poco Phone or OnePlus Nord for reasonable amount of money. But your are right that I'm not the average user since my SIM card is in a Nokia 105, purely because of the battery life.
LG G2 is running a custom ROM with Android 9. And it honestly is usable. Gmail, MS Teams, Puzzle and Dragon, Dualingo and many more applications work without a problem. The probably major difference between how I use the phone and others would probably be that I don't leave applications running in the background. Probably because I think that RAM is the limiting factor here and not the CPU.
I think people don't realise how powerful modern phones actually are, I mean you can record and edit 4k video on today's flagship phones and probably much more.
(ps: yes I carry two phones with me, although I do quite often leave the smart one at home, since the Nokia can handle calls and messages without a problem)
Interesting note, for a 8 years old phone the battery still lasts me the whole day!
The irony of these kind of mitigations is that computers are evolving into C Machines.
Memory tagging only will ship on ARMv9 products in practice, across the board.
The point was more about systems where we have to deal with C, or derived systems languages, and the related culture that performance trumps safety.
I wrote hundreds of thousands of lines of assembler for it and what a fun architecture it was to program for. Enough registers to mean you rarely needed extra storage in a function and conditional instructions to lose those branches and combine that with 3 register arguments to each instruction meant that there was lots of opportunity for optimisation. Plus the (not very RISC but extremely useful) load and store multiple instructions and that made it a joy to work with.
AArch64 is quite nice too but they took a some of the fun out of the instruction set - conditional execution and load and store multiple. They did that for good reason though in order to make faster superscalar processors so I don't blame them!
...At some point I started an x86 on Arm emulator, and managed to run x86 instructions in like 5-7 Arm instructions, without JIT, including reading next instruction and jumping to it - all thanks to the powerful Arm instruction set.
I’ll say the same for the Motorola 68k series CPUs for the exact same reason.
With Intel’s era of total domination nearing an end, I guess we’re seeing some things go full circle.
Similar reasons to why 8 and 16 bit micros stuck around for so long in low cost and low power devices before Cortex M0+ became cheap and frugal enough.
M33s and up might make sense as A64 cores though.
I had to look up how many the ARM1 had in 1985: 25K! 3000nm process.
https://www.astrobe.com/Oberon.htm
https://docs.microej.com/en/latest/PlatformDeveloperGuide/pl...
https://tinygo.org/docs/reference/microcontrollers/trinket-m...
Power consumption is a hard constraint. The benefits of A64 have to outweigh the cost in power consumption, and the benefits just aren't there. Coin cells, small solar collectors and supercaps are among the extremely limited power sources for many Cortex M applications.
Also, worth mentioning that our safety-critical auditor recommends staying on 90nm or larger chips. That’s many of the Cortex chips, and the same as the original iPhone.
Examples—
Z80: Ti graphing calculators
68000: NXP ColdFire
6502: Pacemakers and defibrillators
https://www.elsevier.com/books/computer-organization-and-des...
Who was it that said that R and RISC as a nomenclature was always hogwash. A more apt name would be “load-store architecture”
Onne random thing I remember about LDM/STM was the earliest rev silicon of the Motorola (now Freescale) ARM based DragonBall (then better named the MX line) was the ARM9 core had a bug where LDM/STM would not work with the cache enabled - which of course was horrible, so we hacked gcc to not emit these instructions as a temporary workaround.
hundreds of thousands of lines of assembler ?!?
man...
A72 cores are 10% slower than A73 per clock.
A510 is also supposed to be around 10% slower than A73 -- about the same performance as A72.
The big difference is that A72 is out-of-order while A510 is in-order. This means that A510 won't be vulnerable to spectre, meltdown, or any of the dozens of related vulnerabilities that keep popping up.
For the first time, we can run untrusted code at A72 speeds (fast enough for most things) without worrying that it's stealing data.
Spec pages:
A510: https://developer.arm.com/ip-products/processors/cortex-a/co...
A710: https://developer.arm.com/ip-products/processors/cortex-a/co...
It's speculative execution that causes problems, and in-order CPUs still do speculative execution during branch prediction. It's just that they typically get through less instruction while speculating.
All you need for Spectre is a branch misperception that lasts long enough for a dependent load.
However, it's been over three years since ARM have known about Meltdown and Spectre. There is a good chance the A510 is explicitly designed to be resistant to those exploits.
[0] https://community.arm.com/cfs-file/__key/communityserver-blo... from https://community.arm.com/developer/ip-products/processors/b...
Not sure what the cost of continuing 32 bit support on x86 is for Intel/AMD but does there come a point when Intel/AMD/MS decide that it should be supported via (Rosetta like) emulation rather than in silicon?
Agreed that Arm (esp AArch64) must be a lot simpler.
Then, once you figure out that the instruction at offset 0 is N bytes, ignore the speculative decodings starting at offsets 1 through N-1, and tell the speculative decoding at offset N that it is good to go. That means that it in turn can inform its successors whether they are doing needless work or are good to go, etc.
That effectively means you need more decoders, and, I guess, have to accept a (¿slightly?) longer delay in decoding instructions that are ‘later’ in this cascade.
Well now in 2021 cores are getting wider and wider to have any throughput improvement and power is a bigger concern than ever. Apple M1 has an 8-wide decoder feeding the rest of the extremely wide core. For comparison, Zen 3 has a 4-wide decoder and Ice Lake has 5-wide. We'll see in 3 years if Intel and AMD were just being economical or unimaginative or if they really can't go wider due to the x86 decode complexity. I suppose we'll never really know if they cover for a power hungry decoder with secret sauce elsewhere in the design.
As such, I believe that decoders (even complicated ones like x86) scale at O(n) total work (aka power used) and O(log(n)) for depth (aka: clock cycles of latency).
-------
Obviously, a simpler ISA would allow for simpler decoding. But I don't think that decoders would scale poorly, even if you had to build a complicated parallel-execution engine to discover the start-of-instruction information.
My point is that the ARM's ISA vs x86 ISA is not any kind of asymptotic difference in efficiency that grossly prevents scaling.
Even the simplest decoder on a exactly 32-bit ISA (like POWER9) with no frills would need O(n) scaling. If you do 16-instructions per clock tick, you need 16x individual decoders on every 4-bytes.
Sure, that delivers the results in O(1) instead of O(log(n)) like a Kogge-stone FSM would do, but ARM64 isn't exactly cake to decode either. There's microop / macro-op fusion going on in ARM64 (ex: AESE + AESMC are fused on ARM64 and executed as one uop).
I guess that will affect the constant factor in that O(n) (assuming that’s true. I wouldn’t even dare say I believe that or its converse)
That doesn't change the size or power-requirements of the decoder however. The die-size is related to the total-work done (aka: O(n) die area). And O(n) also describes the power requirements.
If the core is stalled on MMU pages / cache line stalls, then the core idles and uses less power. The die-area used doesn't change (because once fabricated in lithography, the hardware can't change)
> (assuming that’s true. I wouldn’t even dare say I believe that or its converse)
Kogge-stone (https://en.wikipedia.org/wiki/Kogge%E2%80%93Stone_adder) can take *any* associative operation and parallelize it into O(n) work / O(log(n)) depth.
The original application was the Kogge-stone carry-lookahead adders. How do you calculate the "carry bit" in O(log(n)) time, where n is the number of bits? It is clear that a carry bit depends on all 32-bits + 32-bits (!!!), so it almost defies logic to think that you can figure it out in O(log(n)) depth.
You need to think about it a bit, but it definitely works, and is a thing taught in computer architecture classes for computer engineers.
--------
Anyway, if you understand how Kogge-Stone carry lookahead works, then you understand that any associative operation (of which "Carry-bit" calculations is associative). The next step is realizing that "stepping a state machine" is associative.
This is trickier to prove, so I'll just defer to the 1986 article "Data Parallel Algorithms" (http://uenics.evansville.edu/~mr56/ece757/DataParallelAlgori...), page 6 in the PDF / page 1175 in the lower corner.
There, Hillis / Steel prove that FSM / regex parsing is an associative operation, and therefore can be implemented in the Kogge-stone adder (called a "prefix-sum" or "Scan" operation in that paper).
In fixed width ISAs, you have N decoders and all their work is useful.
In byte granularity dynamic width ISAs, you have one (partial) decoder per byte of your instruction window. All but N of their decode results are effectively thrown away.
That's very wasteful.
The only way out I can see is if determining instruction length is extremely cheap. That doesn't really describe x86 though.
The other aspect is that if you want to keep things cheap, all this logic is bound to add at least one pipeline stage (because you keep it cheap by splitting instruction decode into a first partial decode that determines instruction length followed by a full decode in the next stage). Making your pipeline ~10% longer is a big drag on performance.
ARM isn't fixed width anymore. AESE + AESMC macro-op fuse into a singular opcode, for example.
IIRC, high performance ARM cores are macro-op fusing and/or splitting up opcodes into micro-ops. Its not necessarily a 1-to-1 translation from instructions to opcodes anymore for ARM.
But yes, a simpler ISA will have a simpler decoder. But a lot of people seem to think its a huge, possibly asymptoticly huge, advantage.
-----------------
> The only way out I can see is if determining instruction length is extremely cheap. That doesn't really describe x86 though.
I just described a O(n) work / O(log(n)) latency way of determining instruction length.
Proposal: 1. have a "instruction length decoder" on every byte coming in. Because this ONLY determines instruction length and is decidedly not a full decoder, its much cheaper than a real decoder.
2. Once the "instruction length decoders" determine the start-of-instructions, have 4 to 8 complete decoders read instructions starting "magically" in the right spots.
That's the key for my proposal. Sure, the Kogge-stone part is a bit more costly, but focus purely on instruction length and you're pretty much set to have a cheap and simple "full decoder" down the line.
Proof?
> But I don't think that decoders would scale poorly, even if you had to build a complicated parallel-execution engine to discover the start-of-instruction information.
Surely we have empirical proof of that in that Apple's first iteration of a PC chip is noticeably wider than AMD's top of the line offering. On top of this we know that to make X86 run fast you have to include an entirely new layer of cache just to help the decoders out.
We've had this exact discussion on this site before, so I won't start it up again but even if you are right that you can decode in this manner I think empirically we know that the coefficient is not pretty.
What I can say is, I've come up with a decoding algorithm that's clearly O(n) total work and O(log(n)) depth. From there, additional work would be done to discover faster methodologies.
The importance of proving O(n) total work is not in "X should be done in this manner". Its in that "X has asymptotic complexity of at worst case, this value".
I presume that the actual chip designers making decoders are working on something better than what I came up with.
> Proof?
Its not an easy proof. But I think I've given enough information in sibling posts that you can figure it out yourself over the next hour if you really cared.
The Hillis / Steele paper is considered to be a good survey of parallel computing methodologies in the 1980s, and is a good read anyway.
One key idea that helped me figure it out, is that you can run FSM's "backwards". Consider a FSM where "abcd" is being matched. (No match == state 1. State 2 == a was matched. State 3 == a and b was matched. State 4 == abc was matched, State5 == abcd was matched).
You can see rather easily that FSM(a)(b)(c)(d), applied from left-to-right is the normal way FSMs work.
Reverse-FSM(d) == 4 if-and-only if the 4th character is d. Reverse-FSM(other-characters) == initial state in all other cases. Reverse-FSM(d)(c) == 3.
In contrast, we can also run the FSM in the standard forward approach. FSM(a)(b) == 3
Because FSM(a)(b) == 3, and Reverse-FSM(d)(c) == 3, the two states match up and we know that the final state was 5.
As such, we can run the FSM-backwards. By carefully crafting "reverse-FSM" and "forward-FSM" operations, we can run the finite-state-machine in any order.
Does that help understand the concept? The "later" FSM operations being applied on the later-characters are "backwards" operations (figuring out the transition from transitioning from state4 to state 3). While "earlier" FSM operations on the first characters are the traditional forward-stepping of the FSM.
When the FSM operations meet in the middle (aka: FSM(a)(b) meets with reverse-FSM(c)), all you do is compare states: reverse-FSM(c) declares a match if-and-only-if the state was 3. FSM(a)(b) clearly is in state 3, so you can combine the two values together and say that the final state was in fact 4.
-----
Last step: now that you think of forward and reverse FSMs, the "middle" FSMs are also similar. Consider FSM2(b), which processes the 2nd-character: the output is basically 1->3 if-and-only-if the 2nd character is b.
FSM(a) == 1 and ReverseFSM(d)(c) == 3, and we have the middle FSM2 returning (1->3). So we know that the two sides match up.
So we can see that all bytes: the first bytes, the last bytes, and even the middle bytes, can be processed in parallel.
For example, lets take FSM2(b)(c), and process the two middle bytes before we process the beginning or end. We see that the 2nd byte is (b), which means that FSM2(b)(c) creates a 2->4 link (if the state entering FSM2(b)(c) is 2, then the last state is 4).
Now we do FSM(a), and see that we are in state 2. FSM(a) puts us in state 2, and since FSM2(b)(c) has been preprocessed to be 2->4, that puts us at state 4.
So we can really process the FSM in any order. Kogge-Stone gives us a O(n) work + O(log(n)) parallel methodology. Done.
Or under more colloquial terms: you "parse" a potential x86 instruction by "Starting with the left-most byte, read one-byte at a time until you get a complete instruction".
Any grammar that you parse one-byte-at-a-time from left-to-right is a Chomsky Type3 grammar. (In contrast: 5 + 3 * 2 cannot be parsed from left to right: 3*2 needs to be evaluated first before the 5+ part).
On x86, to be able to determine whether a given byte is the start of an instruction, the start of the part of an instruction following the variable-number of prefix bytes or another kind of byte, requires information from the decoding of the previous bytes.
So you may have as many pre-decoders as the fetched bytes, all starting decoding in parallel, but they are not independent and the later pre-decoders need the propagation of signals from the first pre-decoders, so the delay required for determining the initial instruction bytes grows with the number of bytes fetched and examined simultaneously (16 bytes for most x86 CPUs, which on average may contain 4 or 5 instructions).
We do not know for sure where the exact limitation is, but despite the vague and inconsistent information from the vendor documentation, which in a few cases seems to imply better capabilities, until now no Intel or AMD CPU has been shown to be able to decode more than 4 instructions simultaneously (the only exception is that some pairs of instruction are fused, e.g. some combinations of compare-and-branch, and those count as a single instruction for decoding, increasing the apparent number of simultaneously decoded instructions when those pairs occur in the program).
I literally mean in parallel. Computing a FSM from "backwards" and "forwards" simultaneously, as well as "from the middle outward". From all bits, simultaneously, in parallel.
As I stated earlier: Kogge-Stone carry save adder is the first step to understanding this.
Lets take 32-bits of two numbers: A[0:31] and B[0:31]. C [0:32] = A + B.
Simple, yes? Lets focus on purely a singular bit: C[32], the so called "carry bit". It is easy to see that the carry bit (the 33rd bit of the addition operation) depends on all 32-bits of A, as well as all 32-bits of B.
Nonetheless, the carry bit can be computed in parallel with Kogge-Stone (as well as other parallel carry-lookahead adders).
I've explained this process in multiple sibling posts at this point. Please give them a look first, and ask questions after you've read them. I understand that I've written quickly, verbosely, and perhaps inaccurately, but hopefully there's enough there to get started.
But on the other hand verification is a big NRE cost for chips and proving that that second ISA works correctly in all cases is a pretty substantial engineering cost even if the resulting chip is much the same.
These claims are often made and indeed make some sense but is there any actual evidence for them? To really know you'd need access to the RTL of both cutting edge x86 and arm designs to do the analysis to work out what the decoders are actually costing in terms of power and area and whether they tend to produce critical timing paths. You'd also need access to the companies project planning/timesheets to get an estimate of engineering effort for both (and chances are data isn't really tracked at that level of granularity, you'll also need a deep dive of their bug tracking to determine what is decoder related for instance and estimate how much time has been spent on dealing with decoder issues). I suspect Intel/AMD/arm have no interest in making the relevant information publicly available.
You could attempt this analysis without access to RTL but isolating the power cost of the decoder with the silicon alone sounds hard and potentially infeasible.
Likewise, parsing is a well-studied field and parallel parsing has been a huge focus for decades now. If you look around, you can find papers and patents around decoding highly serialized instruction sets (aka x86). The speedups over a naive implementation are huge, but come at the cost of many transistors while still not being as efficient or scalable as parallel parsing of fixed-length instructions. The insistence that parsing compressed, serial streams can be done for free mystifies me.
I believe you can still find where some AMD exec said that they weren't going wider than 4 decoders because the power/performance ratio became much too bad. If decoders weren't a significant cost to their designs (both in transistors and power), you'd expect widening to be a non-issue.
EDIT: here's a link to a die breakdown from AMD
https://forums.anandtech.com/threads/annotated-hi-res-core-d...
After all the relatively simple Thumb extension had been part of the Arm ISA for a long time (and was arguably one of the reasons for its success) and they still decided to go for fixed width.
You'll notice that I understated things significantly. Not only is decoder bigger than the integer ALUs, but it's more than 2x as big if you don't include the uop cache and around 3x as big if you do! It dwarfs almost every other part of the die except caches and the beast that is load/store
https://forums.anandtech.com/threads/annotated-hi-res-core-d...
Original slides.
https://forums.anandtech.com/threads/amds-efforts-involved-i...
The decoder is certainly a reasonable fraction of the core area, though as a fraction of total chip area it's still not too much as other units in core are of similar size or larger size (Floating point/SIMD, branch prediction, load/store, L2) plus all of the uncore stuff (L3 in particular). Really we need a similar die shot of a big arm core to compare its decoder size too. Hotchips 2019 has a highlighted die plot of an N1 with different blocks coloured but sadly it doesn't provide a key as to which block is what colour.
Pretty sure they still have 16-bit support.
It's easier to migrate for a CPU mostly used for Android because it's already had a mix of architectures, there are fewer legacy business applications, and distribution through an app store hides potential confusion from users.
I suppose idly wondering at what point the legacy business applications become so old that the speed penalty of emulation becomes something that almost everyone can live with.
Desktop relies on a lot of legacy support because even today developers make (unjust) assumptions that a certain system quirk will work for the next few years. There's no good reason to drop x86 support and there's good reason to keep it. The PC world can't afford to pull an Apple because there's less of a fan cult around most PC products that helps shift the blame on developers when their favourite games stop working.
Decoders aren't the problem with variable length instructions anyway (and they can be good for performance because they're cache efficient.) The main problem is security because you can obfuscate programs by jumping into the middle of instructions.
For AArch32 not so much.
I expect Qualcomm to stick with Samsung foundry in the next generation, so I am admittedly pessimistic in regards to power improvements in whichever node the next flagship SoCs come in (be it 5LPP or 4LPP). It could well be plausible that we wouldn’t see the full +16% improvement in actual SoCs next year.
https://www.anandtech.com/show/16693/arm-announces-mobile-ar...
It sounds like the thermal issues of the current generation flagship Android chips are expected to remain in place.
v9 has a bunch of neat features that improve security and performance (+ less baggage from ARMv7 which also improves area and power)
Many of them are already optional features in later ARMv8 revisions, however.
v9 brings a bunch of completely new stuff, specially in the security department.
See X1 vs. Firestorm: https://www.anandtech.com/show/16463/snapdragon-888-vs-exyno...
X2 is not even out yet, you have no idea how it may perform.
Why has every post on HN drift into blindly praising an unrelated product?
Some idea how the X2 will perform.
If you read on a bit, there is some question that those performance metrics will be seen in the real world, due to existing thermal issues.
> I am admittedly pessimistic in regards to power improvements in whichever node the next flagship SoCs come in (be it 5LPP or 4LPP). It could well be plausible that we wouldn’t see the full +16% improvement in actual SoCs next year.
Removal of aarch32 has huge implications. And this is not limited to CPU but also touches MMU, caches and more that in v8 had to provide aarch32 compatible interfaces. This lead to really inefficient designs (for example, MMU walk which is extremely critical to performance is more than twice as complicated in ARMv8 compared to v7 and v9).
The space saved by removing these can be used for bigger caches, wider units, better branch prediction and other fun stuff.
Finally, note also that the baseline X1 numbers come from Samsung who we all know are worst in class right now. And they are using an inferior process. Let's see what qcomm, ampere and amazon can do with v9.
By the time products with ARM X2 ship, Apple will be shipping M2 or M3 processors.
ARMv9 tries to unify that a bit [1]. They add a new security model and SVE moves from a side option to a requirement. A bunch of the various v8 instructions also get bundled into the core requirements.
A14 vs X2 is a different question. A14/M1 really dropped the ball by not supporting SVE. I suspect they will support it in their next generation of processors, but the problem is adoption. Everyone is jumping on the M1 units and won't be upgrading their laptops for another few years. As such, they'll be held back for the foreseeable future.
Performance is no question and A14 will be retaining its lead for the next 2-3 years at least. X1 chips are already significantly slower than A14 chips. X2 only increases performance by 10-16% (as pointed out in the article, the 16% is a bit disingenuous as it compares X1 with 4mb cache to the x2 with 8mb of cache while X1 gets a decent speedup with the 8mb version). Furthermore, by the time X2 is actually in devices, Apple will already be launching their own next generation of processors (though I suspect we're about to see them move to small performance increases due to diminishing returns).
[0] https://en.wikipedia.org/wiki/AArch64
[1] https://www.anandtech.com/show/16584/arm-announces-armv9-arc...
Is there any discussion of this online, I'm curious to read more.
A14 is 3.1GHz while A13 is 2.66GHz. A14 is around 20% faster than A13 overall, but it is also clocked 14% higher which gives only around 6% more IPC. Power consumption all but guarantees that they won't be doing that too often.
* only use NEON to save developer time and lose performance and forward compatibility
* only use SVE to save developer time and lose backward compatibility
* Pay to support both and deal with the cost/headaches
Experience shows that AVX took years to adopt (and still isn't used for everything) because SSE2 was "good enough" and the extra costs and loss of backward compatibility weren't worth it.
If SVE were supported out of the gate, then the problem would simply never exist. Don't forget (as stated before) that there's been a huge wave of M1 buyers. People upgraded early either to get the nice features or not be left behind as Apple drops support.
Let's say you have 100M Mac users and an average of 20M are buying new machines any given year (a new machine every 5 years). The 1.5-2-year M1 wave gets 60-70M upgraders in the surge. Now sales are going to decline for a while as the remaining people stick to their update schedule (or hold on to their x86 machines until they die). Now the M2 with SVE only gets 5-10M upgraders. Does it make sense to target such a small group or wait a few years? I suspect there will be a lot of waiting.
SVE vs NEON performance will also hugely depend on the vector length the given algorithm requires, and the stress that the instructions put on the memory subsystem.
Memory hierarchy varies by product and will likely continue to do so regardless of what M2 does.
In the end, I echo eyesee's comment, for best performance one should really use Accelerate.framework on Apple's hardware.
Like other posters mentioned, for vector operations like this, you could be dynamically linking to a library that handles this for you in the best way for the hardware. Then when new instructions become available, you don't have to change any of your code to take advantage.
That said, to get the best performance on vector math it has long been recommended to use Apple's own Accelerate.framework, which has the benefit of enabling use of their proprietary matrix math coprocessor. One can expect the framework to always take maximum advantage of the hardware wherever it runs with no extra development effort required.
With WWDC '21 being only weeks away, I wonder if we're going to see an ARMv9 M2.
The successor to the M1, the lowest-end Apple Silicon running on Macs, is likely expected next year in a rumored redesign of the Macbook Air (whatever the CPU itself is called)
Most likely the M2/A15 will include SVE2 but won't be available for another six months, while the M1+ in the coming Macbook Pro would still be Armv8.5-A.
Perhaps more relevant to the M1, though, is that OS X dropped support for 32-bit apps about 18 months back, with Catalina. The timing there seems just too coincidental to not have been done in order to pave the way for the M1.
This means that you have to do the thunking back and forth yourself to call OS libraries. The only x86 program using this facility so far is Wine/CrossOver on Apple Silicon Macs, which as such runs 32-bit x86 apps...
Judging by some LLVM patches by an apple engineer, they solved this by creating an ILP32 mode for Aarch64 where they use the full 64 bit mode but with 32 bit pointers.
Why run the low power cores in pair? Is this because android _really_ struggles on single core due to the way it was designed? So you basically need two cores even when mostly idle?
I assume if you are doing floating point or SIMD on more than one core then it's time to switch to big cores anyway.
> In terms of the performance and power curve, the new X2 core extends itself ahead of the X1 curve in both metrics. The +16% performance figure in terms of the peak performance points, though it does come at a cost of higher power consumption.
The 16% figure was done with an 8MB Cache compared to a 4MB Cache on X1. i.e In the absolute ( also unrealistic ) best case scenario, with an extra 10% Clock-speed increase, you are looking at single core performance of X2, released with flagship phone in 2022, to be roughly equivalent to A13 used on iPhone 11, released in 2019.
The most interesting part is actually the LITTLE Core A510. I am wondering what it would bring to low cost computing. Unfortunately no Die Size estimate were given.
Cortex-A55 ended up being such a modest, faint improvement over 2012's Cortex-A53. Another point of reference, A53 initially shipped on 28nm.
Given the industry's reliance on ARM's CPU designs, I wonder if it makes sense for prosumers / enthusiasts to keep track of these ARM microarchitectures instead of, say, Qualcomm Snapdragon 888 and Samsung Exynous 2100. Because ultimately the ARM architecture is the cornerstone of performance metrics in any non-Apple smartphone.
Conversely, it'll be interesting to see how smartphone OEMs market their 2022 flagship devices when it's clearly insinuated that the CPU performance on the next-gen ARM cores will not be much of an uplift compared to current gen. Probably more emphasis on camera, display and design.
"The immediate goals for the NUVIA team will be implementing custom CPU cores into laptop-class Snapdragon SoCs running Windows" - Qualcomm’s Keith Kressin, SVP and GM, Edge Cloud and Computing
Perhaps going to smaller technology nodes will make them faster, still? Or is that already part of the prediction?
> The new CPU family marks one of the largest architectural jumps we’ve had in years, as the company is now baselining all three new CPU IPs on Armv9.0.
I believe the deprecation of AArch32 is quite an important change and by itself already warrants issuing a new major version (then there's all mandatory extensions from v8.2+ and SVE2).
They have not claimed to bump the major version on a regular basis from now as you seem to suggest.