RISC-V formal spec public review
github.com
github.com
I wonder how many counterparts to delay slots, stack windows, conditional moves, and other embarrassments we are inadvertently enshrining. There's nothing like hindsight to make you facepalm. (The crypto extension is my bet ATM for most-likely-to-embarrass. But that's without reading it.)
The only way to approach this project sensibly is to assume every single FPGA produced after some near future point will have at least one, and more typically dozens of RISC-V cores scattered around like the multipliers you see in them now, just to try to be competitive.
Personally, I am banking my enthusiasm for when the Bitmanip extension goes in.
Intel and AMD make it work by throwing another 10,000 or 100,000 transistors at it.
It is better to let macro-op fusion hardware identify opportunities to convert a branch-over-move sequence, all by itself. RISC-V is supposed to be all about powerful macro-op fusion.
Clang is really aggressive about generating cmovs. On Gcc you can still use (x & -c) expressions to get nicely pipelined conditional expressions, but Clang stomps them all to cmovs.
Cmov is one of the methods to mitigate Spectre because Intel refuses to speculate loads in them. So, cmov from memory pessimizes your code in cases where you aren't worried what might be sharing your cache.
Sure, this is what makes it good for mispredictable branches - data decompression for instance. If you can't speculate ahead, it's a waste of space in the branch prediction buffer to do it.
But I definitely wish it was a builtin function instead of the compiler generating it, because none of them can guess when to use it.
So, literally exactly the same state that conditional branch instructions depend on? Aside from implementation details (like >Intel refuses to speculate loads in them<, which, to be fair, might be its own problem), "cmovz D S" is just "jnz skip ; mov D S ; skip:" without the useless branch overhead.
There is hope that it doesn't squat a precious branch-prediction slot, hope that it doesn't incur a branch misprediction pipeline stall if it's taken, or not taken, hope that fewer instructions translates to fewer clock cycles, hope that what looks like a copy really just does a register-renaming accounting trick, ...
Or do you mean that the problem is having 3 instead of 2 source operands?
Or that one of the operands is the flag register on x86-like architecture? (but you can just use a normal register being nonzero)
The issue with CMOV is primarily that it's a three operand instruction. If it was a two-instruction operation, it would just be arithmetic, and nobody would care.
A minor, secondary issue is that conditional moves have a data dependency on both of the source operands, unlike conditional branch and move, which means that they are often slower than a predictable branch and move on a fast processor. However, this doesn't matter all that much, because they're optional, so just don't use them when they aren't a good fit.
We did not include special instruction set support for overflow checks on integer arithmetic operations in the base instruction set, as many overflow checks can be cheaply implemented using RISC-V branches. Overflow checking for unsigned addition requires only a single additional branch instruction after the addition: add t0, t1, t2; bltu t0, t1, overflow.
For signed addition, if one operand’s sign is known, overflow checking requires only a single branch after the addition: addi t0, t1, +imm; blt t0, t1, overflow. This covers the common case of addition with an immediate operand.
For general signed addition, three additional instructions after the addition are required, leveraging the observation that the sum should be less than one of the operands if and only if the other operand is negative.
add t0, t1, t2
slti t3, t2, 0
slt t4, t0, t1
bne t3, t4, overflow
In RV64, checks of 32-bit signed additions can be optimized further by comparing the results of ADD and ADDW on the operands.Rust didn't put integer overflow checks in release mode because current CPU don't provide "free" integer overflow detection, RISC V is even a regression over MIPS here..
Not really; the main performance drawback of obligate overflow checks is excessively-constrained semantics of the resulting code making optimization harder, not the direct CPU cost of checking. There are other ways of detecting possible overflows, and Rust does check for overflow in debug mode (this is specifically in order to ensure that most Rust code will work properly even when checks are enabled, and will not simply end up relying on underspecified semantics.)
That said the 'CPU costs' of adding these branch are mostly ICache usage cost so benchmarks aren't very useful..
The problem is that few high level languages support this well either. Having an instruction that traps like division by zero isn't great, because it's quite expensive to turn a machine trap into a high-level exception.
What I would like to try is a separate "addsat" instruction which (a) provides saturating arithmetic and (b) sets a flag if saturation has applied, but does not clear it if it has not been applied. That allows you to write a bunch of arithmetic in a natural way (a1b1 + a2b2 - a3*b3) etc, and only at the end test for overflow. It would behave more like a propagating integer NaN. Since it's a flag rather than a trap a high-level language can more easily convert it to an exception, or provide other processing.
(Saturation arithmetic is useful on its own in some cases, but usually for 8 or 16 bit values)
RISC-V publishes (at least for some cores) the preferred order that you have to emit instructions in order for those to be recognized and fused into efficient micro ops. The compressed (C) extension means that instructions don't take up too much extra space in the I-cache.
Of course we don't know -- because very high performance RISC-V chips don't (yet) exist -- if this will really work when the silicon hits the road, but the plan makes some sense and has the (IMHO) large advantage that it keeps the standard small and easy to implement (in those cases where you don't much care about performance).
There are also extensions, which RISC-V formalizes properly. You can with relative ease add your own addsat instruction through a private extension, and (with somewhat more difficulty) propose a public extension if this is generally useful.
> The Public Review period is: March 29, 2019 through May 13, 2019
Also there is stuff about the carry from 64-bit addition.
(Substitute 16 and 32, for 32-bit units.)
There's no carry flag but I'd say it's still easy to work around that.
; a1:a0 * a2 --> a5:a4:a3
; using t0 as temp reg
mulhu a5, a1, a2
mulu a4, a1, a2
mulhu t0, a0, a2
mulu a3, a0, a2
add a4, a4, t0
; carry a 1 or a 0, overwriting t0
sltu t0, a4, t0
add a5, a5, t0
(edit: switched a1 and a0 for consistency)For a long carry chain, I guess it might look like this:
; Inputs: x = 4-word number, in a3:a2:a1:a0,
; y = 1-word number, in a4.
; Output: x+y, 5-word number, in a4:a3:a2:a1:a0.
add a0, a0, a4
sltu a4, a0, a4 ; a4 = carry bit
add a1, a1, a4
sltu a4, a1, a4 ; a4 = carry bit again
add a2, a2, a4
sltu a4, a2, a4
add a3, a3, a4
sltu a4, a3, a4