This is very a big statement. It's hard for me to think of how they do that, when 8th gen ARM cores are said to be blowing just anything else on size/performance ratio.
Where does SiFive get such an expertise in size optimisation?
This is very a big statement. It's hard for me to think of how they do that, when 8th gen ARM cores are said to be blowing just anything else on size/performance ratio.
Where does SiFive get such an expertise in size optimisation?
It might also help starting from scratch.
I'm also not convinced it's entirely an apples-to-apples comparison. ARM might support more complex instructions that their core don't, and the ARM core might have features like TrustZone.
ARM's 64-bit architecture (AArch64) was also made with similar hindsight, so that's probably not the whole reason.
Most software for Android is not native. My current phone can't run some applications I bought on my Motorola Droid even though I doubt they have one single line of native code. Those applications also don't show up on the Google Store (at least on my phone), so nobody will get them anyway. I had one x86 phone in the meantime, and I didn't see any compatibility issues.
As for the other player in town, Apple, they design everything, silicon, OS and SDK and operate the only application store you can publish to, so, for them, this is also something that can be easily controlled.
And even Apple, who as you said has easily the most control of their ecosystem, is rumored to go to only AArch64 on their next gen chips. They want to move that way, but even they know the issues with moving that direction too fast.
The x86 chip was relying on a process node advantage in order to have a more intelligent memory subsystem that their competitors, to allow them to emulate AArch32 perfoantly compared to their competitors. That process node advantage is now gone and that option isn't available to x86 (and x86 disappeared from the phone market as soon as the writing on the wall was apparent there).
And going back to original point, AArch32 is an albatross around the neck of OoO core design. Features like making nearly every instruction conditional, the restartable load/store multiple instructions, the instruction decoder is almost as complex as x86 (there's almost 1200 instructions in AArch32), instructions can straddle cache line and page boundaries, etc. heavily complicate OoO designs.
Additionally, the one niche that wants powerful cores and isn't dependent on backwards compatiblity (servers), has seen AArch64 only chips.
MIPS, for instance, does not have divide overflow exception. It uses compare, conditional branching and "exception with exception code" instruction.
Most of the time (int32_t/int32_t) division is safe, because divisor is checked for zero in the code logic somewhere else and guaranteed to be non-zero. Sometimes (int64_t/int32_t) it is not safe (higher word is divident can be bigger than divisor) and checking code must be executed in run time. And execution overhead is so little that it is quite good design choice.
You don't have hardware that drains energy constantly for slight slowdown for code that is rare.
As far I can remember, RISC-V does not have divide overflow exception.
As a rule of thumb: if you can put something into software, please do (overflow checks, code scheduling instead of delay slots). Hardware is for things where software can't help (branch prefetch, out of order, etc).
Why? A64 is a completely new design, and unlike RISC-V, by people who've been doing it a while professionally and successfully