ARM’s Cortex A53: Tiny but Important
chipsandcheese.com
chipsandcheese.com
Mobile phones often have background tasks that does not need much CPU power. A53 seems very suitable for this, so it would be nice with some idea of how much power phones saves by using A53 for this instead of a high performance core.
But with software support the model is very effective which is why you see most e-cores these days being relatively beefy OoOE cores that can leave the A53 in the dust. Whether that's Icestorm in the M1, Goldmont on Intel, or A57s on big.LITTLE ARM SoCs.
By compare Intel's been making small but OK revs to the Atom cores. And the new E-core on the new n100 replacement is monstrously faster, yet still small. A potential core-m moment for Intel, a great small chip that expands.
It's a shame, because it was the best design from ARM; they're now focusing on Cortex-A7x and Cortex-X, which aren't anywhere as power efficient[0].
Meanwhile, their revised Cortex-A57 has been surpassed in performance/power/area by several RISC-V microarchitectures, such as SiFive's U74[1], used in the VisionFive2 and Star64, or even the open source XuanTie C910[2][3].
0. https://youtu.be/s0ukXDnWlTY?t=790
1. https://www.sifive.com/cores/u74
A72 in Raspberry 4 is really the pinnacle low power CPU performance if you count $ and compatibility (linux <-> GPU). ~3x can saturate symmetric 1Gb/s with advanced computation and still have calculations left.
You can buy neither, we'll see if the 4 ever comes back. The only thing I'm 100% of is that the 5 will have some drawback compared to both the 2 and the 4.
I can concur the peak of A55; the 5W TDP 22nm RK3566 in my RG353M is mind boggling! HL1 at 60 FPS and soon HL 2 at 30 FPS.
But Risc-V is not progressing with linux, mainly GPU drivers and integration is the problem. Historically the Chinese boards never got any attention and unfortunately you cannot depend on attention happening this time either.
Partly because now companies across the board(ers) are hiding kernel configs again. And "board support packages" are slow to be mainlined if ever.
If you have one hold on to it... I knew it was special!
They are removing 32-bit from modern ARM... but Intel will always have 32-bit!
But now I'm planning a 64-bit linux ARM distro that only runs 32-bit apps!
32-bit is extremely useful. For intel, 16-bit is also very useful. The problem is that there is a strong divide between the people who find 32-bit and 16-bit useful at Intel, and the people who feel a need to force everyone onto UEFI "Class 3+" [1] and that everything else must be squeezed out, shamed, and moved to an emulator.
It sounds like Arm 32-bit support might be around for a while. I mainly wanted to clarify that Intel is dropping their 32-bit and 16-bit support.
Which is a very bold, and probably bad move.
https://www.eenewseurope.com/en/sifive-aims-for-arm-with-hig...
> "The Performance P670 and P470 RISC-V cores are designed to take on ARM’s Cortex A53 and A55 cores, Drew Barbier, senior director of product management at SiFive tells eeNews Europe."
A compare-and-contrast article would make for good reading.
At a similar cost, what’s the real advantage of migrating the current catalog of wearables or IoT products, to RISC-V? There’s a proven and tested platform, widely used in the industry, and the alternative is still trying to catch up.
There are some visual clues. First, the chip pins are labeled in the spec so you can guess they’ll be close to relevant units, and also try to trace their connections throughout the die.
Second, units like memory have an obvious regular structure because they are made from many identical micro-units
Third, if you see, for example, 16 identical adjacent units you could guess this is something that could be dealing with data 16 bits at a time. That narrows it down
There are numerous clues like those.
You could also use tricks like using a thermal camera. What part gets hot when you do certain operations?
Also, in SC5? there was a hallway that had a bunch of plaques from partner companies... and what was awesome was the amount of "Engrish"
Some of the comments were along the same line as "for great justice" (if you are aware of that old meme)
I also personally really value their work—for anyone with intermediate to advanced knowledge of electronics engineering and computers, they are an invaluable source of educational entertainment as traditional mainstream media simply doesn’t cater to such niche audiences.
The software you use plays a rather large role in how the hardware performs. Some people here like to live on the OEM-designed happy path, where things tend to just work. That means using Google Apps for everything, an expectation that the latest video streaming social platforms will open quickly and not stutter, and scrolling the Google Play Store or Google Maps will be a fluid experience.
Others may use simpler apps, or expect less of their phones. I'm in the latter category, and I suspect you are as well. While the BlackBerry KeyOne I use daily was panned by some six months after release in 2017 for being too slow, I instead killed off nearly everything else that would run in the background - including and specifically any Google frameworks and apps.
Some software companies have made a point of taking any hardware gains for granted. Most people have new phones, with fast processors, so some companies will push devs to take shortcuts. I'm quietly indignant about that, though that rant is rather tangental to your original question about how some have such different experiences from yours.
To each, one's own.
But I've always been disappointed by devices that are all A53s.
And when I see devices that have eight A53s and nothing else, I have to assume that they are just trying to trick people into thinking it's a more powerful device than it actually is.
Why would you think that people who actually look up and care about the hardware at the same time are unable to read the first sentence on wikipedia and have no idea what it is? Do you really believe that customers of $100 budget phones are tricked into powerful performance?
I'd guess 25% of customers knew about those at a superficial level, and another 10% actually knew what they should be looking for.
Even an A53 is a super computer when it comes to graphics compared to CPUs of yore.
How much simpler can it be, given that everything seems to already be flat and borderless? As your last sentence alludes to, Windows and other desktop OSs worked perfectly fine with far more complex UIs (including windows) on far less powerful hardware. Mobile UIs seem to be quite primitive in comparison.
In other words, this is entirely a software problem.
E.g. Motorola Moto E30 has a faster processor than G22, but is sold with Go because it only has 2GB RAM and 32GB storage.
It boots Linux in 7 seconds and xfce desktop is pretty snappy.
Kernel is 6.1 and RAM is only 4GB.
Opens Lazarus almost instantly and FPC compiles ARM binaries super fast.
Amazing little machine...
Yes, if the storage is full it can kill both the performance and stability of Android, but devices with slow SoC are slow even with plenty of free space.
We'd find during initial development (i.e., raw, bare Android) that the initial bring up would have good-to-excellent performance, but as the storage began to fill (more "stuff" in the baked-in system/cache partitions, user-installed apps, etc.) it would lag more and more. You'd be surprised how in the early kernels (2.6-3.x series) "iowait"s would slow everything down, UI included, and not just loading speed of apps and such.
News: <https://www.anandtech.com/show/18871/arm-unveils-armv92-mobi...>
Discussion: <https://news.ycombinator.com/item?id=36109916>
This is necessary to make them ISA-compatible with the big cores and medium-size cores with which they are intended to be paired.
Besides the main goal of implementing improved ISA's, they take advantage of the fact that since the time of Cortex-A53 the cost of transistors has diminished a lot and they implement various micro-architectural enhancements that result in a decently greater performance at identical clock frequency, while keeping similar area and power consumption ratios between the small cores like A510 and the medium-size cores like A710, like they are since the first Big.little ARM cores (Cortex-A15 paired with Cortex-A7).
ARM has always avoided to publish any precise numbers for the design goals of the little cores, but it seems that they are usually designed to use an area of about 25% of the area of the medium-size cores and to have a power consumption around 0.5 W per core.
"Generally, when the game is in the foreground, persistent threads such as the game thread and render thread should run on the high-performance large cores, whereas other process and worker threads may be scheduled on smaller cores."
There's also a Wikipedia article [2] which talks a little about scheduling. I imagine Android probably has more specific context it can use as hints to its scheduler about where a thread should be run.
[1] https://developer.android.com/agi/sys-trace/threads-scheduli...
I’m sure Android’s scheduled does things differently but it’s at least an idea of the sort of things which can happen.
For macs (and I assume iOS) the basics are that background processes get scheduled on E cores exclusively, and higher priority processes get scheduled on P cores preferentially but may be scheduled on E cores if the P cores are at full occupancy.
https://www.kernel.org/doc/html/latest/scheduler/sched-capac...
Android does set niceness for processes. See the THREAD_PRIORITY_* constants
https://developer.android.com/reference/android/os/Process
and it uses cgroups too. Process.THREAD_GROUP_* has some Android ones, but different vendors sometimes write their own to try and be clever to increase performance.
Its really quite informative.
NXP has begun to introduce products with Cortex-A55 only recently, but they should always be preferred for any new designs over the legacy products with Cortex-A53, because the Armv8.2-A ISA implemented by Cortex-A55 corrects some serious mistakes of the Cortex-A53 ISA, e.g. the lack of atomic read-modify-write memory accesses.
The people who still choose Cortex-A53 for any new projects are typically clueless about the software implications of their choice.
Unfortunately, there are only 3 companies that offer CPUs for automotive and embedded applications with non-obsolete ARM cores: NVIDIA, Qualcomm and MediaTek. All 3 demand an arm and a leg for their CPUs, so whenever the performance of a Cortex-A55 is not enough it is much cheaper to use Intel Atom CPUs than to use more recent ARM cores, e.g. Cortex-A78.
> non-toxic IP cores with enough performance from RISC-V
Why would anyone also their designs to be used for free or for cheaper than ARM does?
I don’t see how can high-end/competitive RISC-V cores could be fully open/free and without that how is it better than ARM.
arm and x86 are not royalty free ISAs, RISC-V is, then RISC-V is mechanically a better choice, until it does a good enough job.
Only ARM can license ARM cores to others.
Using RISC-V, any company who can design their own cores can also offer them for licensing.
There's already tens of companies offering hundreds of cores for licensing.
This is much better than ISA-enforced vendor lock-in, which is the situation with x86 and ARM.
And the companies that don’t want to make it their core business but can afford enough resources (e.g. Google, Apple, Amazon) would just use them to leverage their core products.
I could only see this on the lower end where margins/required R&D investment are relatively low.
Ofc, worldwide royalty free is not enough (or just be allowed to implement the ISA...), silicium is really about performance, and I sincerely hope RISC-V will end up providing microarchitectures (open or not) "good enough" to do the job.
I am perfectly aware RISC-V will fail if not providing at scale really good implementations. Rumors say "really performant" implementations are not expected before 2024.
The problem with this argument is that it ignores the cost of creating the microarchitecture. It’s almost certainly cheaper to license an Arm A series core than to create a comparable RISC-V core from scratch.
Sure we have a firm like SiFive that licenses RV cores to third parties but the existence of firms like Arm and SFive shouldn’t be taken for granted. If rumour has it SiFive were almost taken over by Intel. Thankfully Nvidia were stopped from buying Arm.
If you wish for Arm’s demise you may get the end of their business model and that probably isn’t a great outcome.
But I agree that without _REALLY_ performant implementations, risc-v WILL fail and rumors say that things won't start to get serious before 2024. Nevertheless, in its current state, failure is still a more than valid outcome.
What is interesting imho is that they WILL get serious by 2024, including P670, Veyron, Ascalon. All of them seem to implement RVA22 + ratified V.
This is a very fast timeline for RISC-V, which privileged spec was only ratified in 2019.
It all has been happening much faster than even the most optimistic estimates were.
As far as I know and specs wise, RISC-V has been kind of ready for a while, only missing very high performance implementations.
My first real risc-v target is a 100% RV64 assembly keyboard firmware though. Looking at mango pi pro mq boards, but I wish we had 'smaller' RV64 GPIO/USB boards for that, maybe a small GPIO/USB board with a USB block+FPGA with enough gates to instance a small RV64 core.
These are known "detonators" set up for next year. There's more we don't even know of. It is going to happen.
RISC-V is inevitable.
>As far as I know and specs wise
December 2021's batch of extensions done the magic. No hardware out there implements them.
Once it arrives (e.g. Veyron and Ascalon), the thorough disruption of the whole CPU landscape, from microcontrollers to supercomputers, starts.
The acceleration and massive momentum that we've seen to date is nothing compared to what's to come.
I want to share you optimism, but I advise you to keep your cool. There is a long road to reach the performance of x86/arm microarchitectures (it is harder for risc-v since the "market is over saturated"). And those performant implementations must get access to the best silicium node process... and that...
Veyron's supposed to have actual server boards on sale late this year.
Ascalon, CPU chiplets and a full AI accelerator product using them, succeeding their current one, which iirc uses SiFive IP.