New SiFive RISC-V core P650 with 40% IPC increase
sifive.com
sifive.com
I expect a lot of tape-outs to happen this month, as core vendors were probably holding for the announced ratifications, in fear of last minute changes. Next year is going to be exciting.
[0]: https://riscv.org/announcements/2021/12/riscv-ratifies-15-ne...
That being said, I don't think it's the worse thing in the world like some do. The focus now should be on compiled code since JITs by definition can make runtime descions on if some future extension that fixes this deficiency exists or not. The J extension has stalled for the moment, but with these other extensions ratified there should be more bandwidth available hopefully.
Or is that what we are trying to get away from?
The complaint is valid, IMO, and would show up on the filtration test they used to come up with ops if they were working with JITs too rather than just what's in AOT code.
Hence why I don't think "RISC is the future" unlike a lot of other proponents; I think a CISC with uop-based decoding will be more scalable and performant. Even ARMs have moved a little in that direction.
The riscv stans keep saying that, but nobody has given a demo or shown benchmarks afaik, even under simulation. So it's just handwaving.
It's not only javascript, of course. int overflow in C is an error condition (undefined behaviour) that compilers usually don't try to trap (the -trapv option in gcc and clang enables trapping at some performance cost, so it's rarely used and we get continuing bugs and vulnerabilities as a result. Ada mandates trapping unless you enable an unsafe optimization which is, um, enabled by default in GNAT). Riscv increases that performance cost considerably from what I can tell. That's the opposite of what we needed.
I'm no CPU architect but I know they are able to signal overflow in floating point arithmetic, since IEEE 754 requires that. So I don't understand why they can't do it for integers.
Background reading: https://huonw.github.io/blog/2016/04/myths-and-legends-about...
If you only want safety, then trapping or not, signalling or not does not matter at all. It is UB that causes safety problems, not the overflow itself. And RISC-V mandates the overflow handling manner. No UBs.
Throw on arithmetic overflow is a language choice. And at least Rust thinks that arithmetic exception everywhere is not necessary for security.
The only related problem with no overflow trapping is that dynamically typed languages needs numerical type conversion on overflow. But TBH, if a numerical javascript program often generates 1.7E308, then it's a terrible program that no one should care.
MIPS is kind of the spiritual ancestor to RISC-V
I am curious about the final design. Would be interesting to hear how people think it compares with ARMs scalable vector extensions.
I like it. It's fairly simple and clean, yet powerful.
There was also some discussion here in HN months ago, about an article comparing RISC-V V extension and ARM SVE.
The article itself got several things wrong about V, but the discussion[0] was interesting.
This will be beautiful to watch.
If Apple bought all future 3nm capacity from TSMC, good luck trying to compete.
Are you assuming the competition will just sit and do nothing?
There is a lot of wishful thinking that using RISC-V magically makes all of the SOC more open. It doesn’t!
I agree, this is wishful thinking.
such as the $30 Sparkfun Red or the $20 Lofive boards. Those are for running an RTOS, not Linux, but they compete with Arduino, mbed, teensy, and other ARM Cortex M series microcontrollers.
A price target of $10 is something you'll only hit with massive scale-up.
Those are not computers
That seems to imply a certain integer arithmetic performance, but I wonder what the floating point performance is. They could have just said "X flops".
Comparing to other benchmarks at [1], I have no idea, because they all have denormalized results, so totals, rather than per GHz per core. Nice reporting.
How fast is this thing? Pentium? first gen i3? current gent ryzen 5? The fact that they are being so obtuse about it leads me to believe performance isn't great.
[1] https://www.spec.org/cgi-bin/osgresults?conf=cint2006;op=dum...
Curious how IP pricing compares to ARM in this case and how much would I need to put on top of it to tape out own batch of processors
There's several vendors besides RISC-V offering cores for licensing. There's even some OSHW cores that can be freely used.
Even if we choose to ignore the technical prowess of being a true 5th generation RISC ISA built with hindsight no other ISA has, what's IMHO a big deal in RISC-V is the mere availability of this market of cores.
It poses a threat to ARM's business model, where ARM licenses cores and ISA, but nobody else than ARM can license cores to others.
Why do all the riscv fans Conveniently ignore aarch64 when they make statements like this? It was in fact a completely clean new design, based on hindsight, by people who know what they are doing, and with no legacy Cruft.
The "by people who know what they are doing" thing is just pure FUD. Sure, ARM employs some competent people, but no more so than IBM, Intel, AMD or the various members of RISC-V International.
I think however that RISC-V isn't that much worse and because of the freedom we will almost certainly see more implementation of RISC-V. I'd be watching Tenstorrent, SiFive, Rivos, Esperanto, and maybe Alibaba/T-Head.
aarch64 seems poorly designed to me.
ARMv7 had thumb, but for some reason ARMv8 did not incorporate any lessons from that. As a result, code density is bad; ARMv8 binaries are huge.
ARMv9, to be available in chips next year, is just a higher profile of required extensions, and does nothing to fix that.
Ever wonder why M1 needs such huge L1 cache? Well, now you know.
Considering ARMv9 will be competing against RVA22, I don't have much hope for ARM.
RISC-V has only one good feature for code density, the combined compare-and-branch instructions, but even this feature was designed poorly, because it does not have all the kinds of compare-and-branch that are needed, e.g. if you want safe code that checks for overflows, the number of required instructions and the code size explode. Only unsafe code, without run-time checks, can have an acceptable size in RISC-V.
ARMv8 has an adequate unused space in the branch opcode map, where combined compare-and-branch instructions could be added, and with a larger branch offset range than in RISC-V, in which case the code size advantage of ARMv8 vs. RISC-V would increase significantly.
While the combined compare-and-branch of RISC-V are good for code density, because branches are very frequent, the rest of the ISA is bad and the worst is the lack of indexed addressing, which frequently requires 2 RISC-V instructions instead of 1 ARM instruction.
Many things could be said about ARMv8, but that it has good code size is not one of it. It does, in fact, have abysmal code density. Both RISC-V and x86-64 produce significantly smaller binaries. For RISC-V, we're talking about a 20% reduction of size.
There's a wealth of papers on this, but you can verify this trivially yourself, by either compiling binaries for different architectures from the same sources, or comparing binaries in Linux distributions that support RISC-V and ARM.
>where combined compare-and-branch instructions could be added, and with a larger branch offset range than in RISC-V
If your argument is that ARMv8 could get better over time, I hate to be the bearer of bad news. ARMv9 code density isn't any better.
>and the worst is the lack of indexed addressing, which frequently requires 2 RISC-V instructions instead of 1 ARM instruction.
These patterns are standardized, and they become one instruction after fusion.
RISC-V, unlike the previous generation of ISAs, was thoroughly designed with hindsight on fusion. The simplest microarchitectures can of course omit it altogether, but the cost of fusion in RISC-V is low; I have seen it quoted at 400 gates.
The one fusion implementation I'm aware of if the SiFive 7-series combining a conditional branch that jumps forward over exactly one instruction. It turns the instruction pair into predicated execution.
I agree with everything else. In particular the code density. Anyone can download Ubuntu or Fedora images for the same release for amd64, arm64, and riscv64. Mount them and run "size" on any selection of binaries you want. The RISC-V ones are consistently and significantly smaller than the other two, with arm64 the biggest.
I've heard of that feature before somewhere else. It gave the company that invented it unparalleled code density in their 32 bit systems and propelled them to the heights of success in mobile devices. What was their name? Wait .. oh, yes ... ARM.
Why they forgot this in their 64 bit ISA is a mystery. The best theory I can come up with is that they thought the industry had shaken out and amd64 was the only competition they were going to have, ever. Aarch64 does indeed have very good code density for a fixed-length 32 bit opcode ISA, and comes very close to matching amd64. They may have thought that was going to be good enough.
Note: the RISC-V "C" extension is technically optional, but the only CPU cores I know of that don't implement it are academic toys, student projects, and tiny cores for use in FPGAs where they are running programs with only a few hundred instructions in them. Once you get over even maybe 1 KB of code it's cheaper in resources to implement "C" than to provide more program storage.
There are 2 ways of programming a loop that addresses memory with a minimum of instructions.
One way, which is preferable e.g. on Intel/AMD, is to reuse the loop counter as the index into the data structure that is accessed, so each load/store needs a base register + index register addressing, which is missing in RISC-V.
The second way, which is preferable e.g. on POWER and which is also available on ARM, is to use an addressing mode with auto-update, where the offset used in loads or stores is added into the base register. This is also missing in RISC-V.
Because none of the 2 methods works in RISC-V with a minimum number of instructions, like in all other CPUs, all such loops, which are very frequent, need pairs of instructions in RISC-V, corresponding to single instructions in the other CPUs.
Also, high performance Aarch64 and POWER implementations are likely to be splitting those instructions into two decoupled uops in the back end.
Performance-critical loops are unrolled on all ISAs to minimise loop control overhead and also to allow scheduling instructions to allow for the several cycle latency of loads from even L1 cache. When you do that, indexed addressing and auto-update addressing are still doing both operations for every load or store which, as well as being a lot of operations, introduces sequential dependency between the instructions. The RISC-V way allows the use of simple load/store with offset -- all of which are independent of each other -- with one merged update of each pointer at the end of the loop. POWER and Aarch64 compilers for high performance microarchitectures use the RISC-V structure for unrolled loops anyway.
So indexed addressing and auto-update addressing give no advantage for code size, and don't help performance at the high end.
not true
you could compile and compare, say, gcc (cc1) of the same version on arm64 and rv. arm64 has larger binaries
I used to think so too, until I asked some more knowledgeable people about it. Turns out the lesson IS that not having it is better. Fixed-sized instructions make a decoding significantly simpler, making it much easier to make very wide front ends
You can put as many of these modules side by side as you want. There is a serial dependency between them in that each block has to tell the next block whether its last 16 bits are the start of a misaligned 32 bit instruction or not. That could become an issue with really really wide but for something decoding e.g. 16 bytes at a time (4 to 8 instructions) it's not an issue.
There is a trade-off between a little bit of decoder complexity and a lot of improved code density -- but nowhere near to the same extent as say x86.
I had thought that Thumb 1 had serious shortcomings, which is why they ended up needing Thumb 2.
I'm not sure I follow this, but it reminds me to ask: does RISC-V allow for designs to have both efficiency & performance cores like the ARM big.LITTLE concept? Has anyone made one yet?
To date the examples of this that have been shipped to the public have used cores with similar microarchitecture, but a different set of extensions.
For example the U54-MC in the HiFive Unleashed and in the Microsemi Polarfire SoC FPGAs use four U54 cores plus one E51 core for "real time" tasks. The E51 doesn't have an FPU or MMU or Supervisor mode. The U74-MC in the HiFive Unmatched is similar.
Alibaba's ICE SoC, which you may have seen videos of running Android, has two C910 Out-of-Order cores (similar to ARM A72/A73) implementing RV64GC, and a third C910 core that also has a vector processing unit with two pipes with 256 bit vector ALU each, plus 128 bit vector load and store pipes.
https://www.extremetech.com/wp-content/uploads/2018/07/arm-r...
It will be a PR disaster long remembered. One for the textbooks.
No amount of FUD will save ARM. Only pivoting into a different business model could.
Given M1, Graviton etc etc that’s a bold statement.
x86-64 is much worse than ARM. It's a literal clusterfuck. And yet.
A high performance implementation of ARM, which is a much better ISA than x86-64, was something expected to happen sooner or later. It did not surprise me.
11+ SPECInt2006/GHz is comparable to Apple Icestorm microarchitecture. Apple Firestorm microarchitecture is roughly 2x better at 22 SPECInt2006/GHz.
M1's L1 cache is huge, as a workaround to ARMv8's poor code density. Larger cache means lower clocks, unfortunately there's no way around speed of light.
What? Where is this claim coming from?
Obviously you can't even think about comparing it further with Intel & AMD, but when you look at the history of something like ARM(which i believe is 30-40 years old), riscv came a long way pretty fast, and the good thing it's a solid choice for the future due being open.
I guess the RISC-V will conquer the desktop the same year Linux will.
Note, that switching to ARM from x86 is a pain, esp if you depend on proprietary software.