An ex-ARM engineer critiques RISC-V
gist.github.com
gist.github.com
RISC-V means three instructions instead of two in the best case. It requires five or more instead of two in bad cases. That’s extremely annoying since these code sequences will be emitted frequently if that’s how all math in the language works.
Also worth noting, since this always comes up, that these things are super hard for a compiler to optimize away. JSC tries very aggressively but only succeeds a minority of the time (we have a backwards abstract interpreter based on how values are used, a forward interpreter that uses a simplified octagon domain to prove integer ranges, and a bunch of other things - and it’s not enough). So, even with very aggressive compilers you will often emit sequences to check overflow. It’s ideal if this is just a branch on a flag because this maximizes density. Density is especially important in JITs; more density means not just better perf but better memory usage since JITed instructions use dirty memory.
Overflow is part of the result, so maybe include extra bits to each register that can be arithmetic destination. These bits are not included in moves, but could be tested with new instructions.
Another way that avoids flags is new arithmetic instructions: add but jump on overflow. Maybe this is reduced to add and skip next instruction except for overflow, but maybe things are simplified if the only allowed next instruction is a jump, so the result is a single longer instruction.
The best design is part of the core ISA and not and extension since overflow checking is fundamental to modern languages.
If you doubt that dynamic languages are using overflow checking in the way that I describe then it’s not because of lack of papers on the subject.
If you did overflow checks rarely then what you say is a very good point indeed. The key thing is just the frequency of this stuff in modern languages.
Basically, if there's a bottleneck to x86 code, Intel has run into it, profiled for it and generally optimized around it both in their microarchitectures and in the their C compiler.
I like the Mill CPU approach, where every "register" (it doesn't have named registers actually) has the full set of status bits associated with it, and not just for overflow. Things like "not a result" (NaR), which can represent the result of a failed speculative load for example (because the process doesn't have permission to read from that page, for example).
The status bits part in general, or the speculative load stuff?
They allegedly have all this working, privately. They haven't released any development tools or such to the public.
I've often toyed with the idea of writing an instruction-level simulator (as opposed to the RTL sim or whatever they have internally). But even sticking to the public information, I'd likely be infringing on their patents.
"Several ISA features, including implicit condition codes and predicated moves, are onerous to implement in aggressive microarchitectures. Yet, their complexity often does not result in higher performance because their semantics were ill-conceived. For example, x86 provides a conditional load instruction, but, if the unconditional load were to cause an exception, it is implementation-defined whether the conditional version would do so. Thus, a compiler can only rarely use this instruction to perform the if-conversion optimization.
Recognizing the inefficiency of their conditional operations, Intel’s recent implementations go to some lengths to fuse comparison instructions and branch instructions into internal compare-and-branch operations."
Yeah, doing condition codes right is complicated but then that's the purpose of microarchitecture, to factor that complexity out.
RISC-V succeeds in having a minimalist to a fault design which is appropriate for a certain design point. ARMv8 and x86_64 are more useful for a broader set of designs. The burden now is on RISC-V to show that their minimalist approach is fast+efficient rather than just simple. You have to get something out of the simplicity; otherwise what you get is design debt.
add t2, t1, t0
sov t3, t1, t0
bnez t3, overflow
Why this way? "extra bits on destination registers"- this is really flags. The flags have to be preserved during interrupts, so extending the registers is not so easy (I think it just reduces to classic flags)."add but jump on overflow" or "add and skip on no overflow"- I don't like this because you can not break it into separate operations without flags. I think you might have to add hidden flags in a real implementation.
An add followed by an sov could be fused, but requires an expensive multi-register write. Fusing maybe could be more likely if the destination is always to a fixed destination register:
add t2, t1, t0
sov tflags, t1, t0
bnez tflags, overflowaddsov t2, t1, t0
I don’t think these folks have seen what modern languages do.
ISO C is therefore a three-way compromise between three communities: Software authors, compiler authors, and CPU manufacturers. You will naturally end up with some decisions that dissatisfy some members of each.
1) it’s pretty important to have easy syntax for math with overflow checks. Otherwise there will be bugs. Bad bugs. We have those and it sucks.
2) it’s pretty important that if the overflow wraps in ISA then it wraps in language semantics but C just says “meh whatever” for signed overflow. That leads to even more bugs.
GCC has intrinsics for integer math with overflow checks
How much business code out there burns most of its cycles in Java/C#/Rust/Go? Having even one of those (lets face it: Java) would have gone a very long way. How much client computing spends most of its cycles in (JITted) javascript?
I understand the tradeoff that drives this omission, though. How many VC megabucks (Tens? Hundreds?) would they have had to spend commercially on a bet? Alternatively, how many grad student-years grinding through low payoff work to prove yea/nea on the value of checked arithmetic?
Lets say that you make a hip-shot bet on some checked arithmetic support without the supporting toolchains and application code to back up the design choice. You could end up making a different set of mistakes instead and end up in the same situation.
I think the only mistake was in finalizing the ISA without any support for checked arithmetic. My belief is that doing it well will not be orthogonal to the rest of the ISA's design, and therefore is a poor candidate for an extension.
Note macro op fusion is widely used for other architectures already, particularly ones like x86 where what the processor actually runs looks nothing like the machine code.
It doesn’t matter if they’re fused or not if the reduced instruction density increases memory usage and puts more pressure on I$.
Also, I don’t buy the whole fusion argument on the grounds that having to fuse super complex (5 instruction or more) sequences adds enough complexity that you’ve got opportunity cost. Much better for everyone if the CPU doesn’t have to do that fusion. That’s the whole point of good ISA design - to prevent the need for fusing in cases you’re doing something super common.
Edit: Here's the thesis about the design decisions in the C encoding: https://people.eecs.berkeley.edu/~krste/papers/waterman-ms.p... See also the diagram on page 62 of this document: https://people.eecs.berkeley.edu/~krste/papers/EECS-2016-1.p...
The thing I see is this: you can also add compressed encodings for any ISA. And that has its own costs (it’s harder to decode and it’s harder on software that wants to do bidirectional analysis of machine code). So “my isa has shortcomings but it’s cool because compression” isn’t a perfect argument since if your isa lacks those shortcomings then you still benefit from compression and you don’t need it as much, which is better.
RISC-V compact instructions don't require special modes and run in fully-mixed mode with 32-bit instructions without all the penalties thumb has (they are literally just extended into their 32-bit counterparts internally).
High code density became less valuable, was the point.
(IIRC they removed that one on x86-64)
not necessarily, since
1. the compressed versions are basically 1-1 to the non-compressed versions
2. the uncompressed versions are more for study/academia and clarity; its expected that irl only compressed instructions are used (for instructions that can be compressed)
more info: https://riscv.org/wp-content/uploads/2015/11/riscv-compresse...
Can you give an example of someone advocating for 5 instruction fusion? Normally it's limited to three.
You need fast sequences for all of these variants:
- add or sub or mul
- signed or unsigned
- 32 bit or 64 bit
Some of those need 5 instructions. I don’t remember which adventure you need to pick to get 5.
sadd32, sadd64, uadd32, uadd64, ssub32, ssub64, usub32, usub64, smul32, smul64, umul32, umul64
I don't think it's a fusion problem; even if you did fuse these sequences, they'd still be bad, since they'd be writing lots of extra registers.
Bitfield insertion is only one instruction in most RISC ISAs, but 5 or more in RISC-V.
That said, until it gets ratified by the consortium and implemented in silicon its still just a (well-researched) wishlist.
add t0, t1, t2
slti t3, t2, 0
slt t4, t0, t1
bne t3, t4, overflow
So two extra registers needed also..edit: so I now think a good extension to add overflow checking to RISC-V is with an instruction that works like "slt"- call it "sov", set if add would overflow:
add t0, t1, t2
sov t3, t1, t2
bnez t3, overflow
add/sov could be fused..The only way to know there isn't a dependency is if that register gets clobbered by something else very soon afterwards.
But this whole topic of "checking for signed overflow is expensive" is overblown. It's simply not that important an operation, especially in the context of those languages that do it a lot also doing a lot of memory references, which are far more expensive.
Adding arbitrary completely unknown integers is pretty rare. If you know both numbers are greater than zero then a single compare-and-branch is all you need. If one of the numbers is a constant then a single compare-and-branch is all you need.
auipc lr, zero, .LONG_TARGET20
jalr lr, lr, .LONG_TARGET12
Only one register (the link register) is clobbered, so the pair can be fused into a single wide jump-and-link.So in parent's example, sov might fuse with the following bnez, but it likely wouldn't fuse with the preceding addition.
Quite a few things determine whether a fusion is doable or not. In addition to the number of destination registers, you do, to a more relaxed extent, care about source operands, but also things like ‘does this fit nicely in a single pass through the pipeline?’ and even just ‘is this materially beneficial?’
Lots of cores (but not all) can write two registers from a fused instruction, given the right conditions, and sov does rerun the addition, so add-sov fusion sounds very doable to me.
(Quoted from OPs comment)
This isn’t a subject I’m an expert at but wouldn’t this mean there’s already some sort of translation going on already on the other systems? So it’s mostly just added end user work not a giant performance loss on RISC?
It would just then simply be a layer of abstraction that is lost.
The x86 front end breaks down big instructions into smaller RISC-like micro-ops and then fuses/re-orders/optimizes/etc and runs those instead. There’s pros and cons, the con being sheer complexity and power budget, the pros are that it’s an abstraction so the microarchitecture can change completely without recompiling your code —- and you get CPU specific optimizations too. The CPU is basically emulating x86.
You could in theory build an x86 CPU with a RISC-V core behind that decoder.
This is a widely held meme, but the internet at large doesn't have any evidence to back it up. A couple of publicly visible engineers that do have experience are on record as saying that cell-phone-class competitive x86 was absolutely possible. Intel and AMD chose not to pursue those markets.
The expensive parts of a high-end CPU aren't normally in the instruction decode part. They are in the branch prediction, branch mispredict recovery, forwarding networks, memory re-ordering, and so on. Anything short of a dataflow ISA has little impact on those structures.
Weren't there Windows Phone devices with x86 SoC, but they weren't competitive?
Commercially made x86 Android phones exist, the most popular to my knowledge were some of the Asus ZenPhone models.
When Atom was released in 2008, the A9 had already been announced (a year before). A9 was around 10-15% faster per clock and were often multicore meaning a 1.5GHz chip was faster in all metrics over the 1.6GHz Atom.
A few articles came out June/July of this year with Analysts saying that Intel had spent over 10 Billion dollars trying to break into the mobile market with no success. ARM's current R&D budget (according to Nvidia a month or so ago) is 0.5B. If they spent that much every year since 2000, they would barely match Intel, but their entire market cap in stayed under 4B all the way until 2009. Remember, that R&D includes their high-end ARM cores, but also GPU, midrange designs, various microcontroller designs, a realtime OS, ARM tooling, NPUs, various kernel support, etc
If that much money can't fix up x86 to keep up with a budget a fraction of the size, I take that as proof that the ISA really does matter.
[1] https://www.usenix.org/system/files/conference/cooldc16/cool...
I don't think you can generalize from that result to much of anything.
That struck me as one of the few apples to apples comparisons ever of the instruction sets at the high end, from a party not really incentivized to bend the truth one way or another.
But it could have easily come from something like the relaxed memory model, or they could have just been overly optimistic. The chip was cut after all.
With the exception of Atom, all recent Intel designs have been a RISC core with a CISC decoder slapped on top. Everything else being equal, the simpler decoder will create a smaller chip. Because the decoder is always running all-out, the simpler decoder will also use less power.
x86 instructions are multi-length from 1 to 15 bytes. The cost to slice that up is always going to be bigger than fixed-length instructions. RISC-V has variable length in theory, but in practice, compact instructions extend into 32-bit instructions with some bits added which is important for decode cache while longer instructions are ignored.
Because there's a maximum tolerable decode latency of a very few cycles and latency increases with cache size, decode cache size has a definite cap. x86 has a couple orders of magnitude more potential instructions than RISC-V. More instructions translates into a lower hit rate for the same size cache barring any heuristics (more on that below).
Matching variable-length arrays to an unknown set of arrays in cache is inherently a hard problem. Every solution has tradeoffs and the resulting heuristics are bad for computing (see below). In contrast, searching for a match on a fixed 32-bit array has much more simple general solutions that don't require tradeoffs.
C and CPU designs feed off each other. Let's say there are instructions X and Y which can do equivalent things. x86 engineers played around with both and got a lucky insight into how to make X a bit faster. Compiler writers jump on it and start using the faster solution. x86 engineers now all but stop looking to improve Y and spend their time tinkering with X instead. Compilers now focus even more heavily around not just X, but any instructions more closely associated with X.
In that entire (true) story, nobody gave a second thought to whether the final result of Y would have been faster overall if not for the lucky break with X. If x86 developers were actually free to choose whichever instructions they wanted, x86 decode would be much, much slower than it appears to be. This self limitation argues that perhaps a more RISC-like ISA is inevitable.
A new ISA where everything is used would definitely have complexity, transistor, and power disadvantages vs a new ISA that didn't make that mistake.
Don't think its strictly accurate to say that Intel 'chose not to pursue' the mobile SoC market - IIRC they tried, made little progress and gave up having spent a lot of money in the process.
Can't wait to see more information about the goldmont microcode work to see if that holds for intel as well as it does for AMD.
// &(array[offset])
slli rd, rs1, {1,2,3}
add rd, rd, rs2
The sequence in the article uses what Intel calls the fast case but it still wouldn't qualify for Berkeley's two instruction fusion. Dunno if anyone does three instruction fusion.As an aside, LEA is never getting added to the base RISC-V nor should it be. But I'm surprised it isn't considered for an extension.
[1] https://www2.eecs.berkeley.edu/Pubs/TechRpts/2016/EECS-2016-...
I'm not sure that is the idea, given that RISC-V is targeted at processors so low-end that they don't even implement multiply.
Off-topic, but could you point me to more details on this? Someone (else?) recently mentioned octagon analysis in JSC in a HN thread. I grepped through the sources at the time but didn't find any indication that it exists. At least not under the name "octagon".
Look at DFGIntegerRangeOptimizationPhase.cpp
RISC-V has some closely-related sharp corners in indexed address arithmetic as well. Some choices for the type of the index variable perform much worse on rv64.
Consider: an LP64 machine uses 32-bit integers for 'int' and 'unsigned', but 64-bit integers for `long`, `size_t`, `ptrdiff_t` and so on.
If you use an array index variable of type `unsigned`, then the compiler must prove that wraparound doesn't happen. That's pretty weird considering that half the point of using unsigned is to elide such proofs of correctness. If it cannot prove the absence of unsigned wraparound, then it will be forced to emit zero-extension sequences prior to using the index variable to generate the addresses.
ARMv8 side-steps the whole problem by providing indexed memory addressing modes that include the complete suite of zero and sign extension of a narrow-width index in the load or store instruction itself.
So here we have an example of a three-way system engineering choice.
- Provide a small amount of hardware that performs the operation on-demand.
- Provide new and inventive forms of value-range analysis in the compiler. Despite decades of research into this problem, the world's best solutions still frequently saturate at "the entire width of the type the programmer requested".
- Change the habits of the world's C programmers.
RISC-V chose options 2 and 3.This is usually why your array indexing should be done with an iterator or size_t :)
The problem isn't with unsigned types generally. Its with subregister unsigned types. So, size_t and uintptr_t are fine. uint32_t, uint16_t, uint8_t (on LP64 ABIs) are pessimized and demand zero-extension instructions (or proofs that they can be safely elided) prior to causing side-effects. uint64_t on a LP128 ABI would also be problematic.
signed 32-bit int is also fine... because RISC-V specifically has a suite of arithmetic instructions that unconditionally sign-extend from bit 31. Even without those, it would still be fine because the carve-out for undefined behavior is wide enough for INT_MAX+1 to remain positive. Same thing for all of the other narrow-width signed integer types. If you increment SHRT_MAX and then use it to index memory, its perfectly legal undefined behavior to access base + SHRT_MAX+1 instead of base + SHRT_MIN.
However, that's not legal for the unsigned types. They are all mandated to wrap in 2's complement. base + UINT_MAX+1 must access base + 0 when the index is `unsigned int`, even on a 64-bit machine.
Ironically, given the topic, that's actually not true, because `base + UINT_MAX+1` is `(base + UINT_MAX)+1` (with a pointer, not a unsigned int, as the temporary value). That should probably be `base + (UINT_MAX+1)`.
But isn't size_t (or ptrdiff_t) the preferred indexing type in C for this reason (among others)? Sometimes you of course do want wrap around modulo semantics but that's much rarer, right?
I think an important design point here is that the languages that need a lot of dynamic overflow checks are primarily used on beefier CPUs so if you can get around the code size issue, making it performant only on more capable designs is fine since the overflow check will be rare on simpler CPUs.
1. Folks totally run JS and other crazy on small CPUs.
2. Other safe languages (rust and swift I think?) also use overflow checks. It’s probably a good thing if those languages get used more on small cpus.
3. The C code that normally runs on small cpus is hella vulnerable today and probably for a long time to come. Compiling with sanitizer flags that turn on overflow checks is a valuable (and oft requested) mitigation. So theres a future where most arithmetic is checked on all cpus and with all languages.
And yeah, it’s true that the overflow check is well predicted. And yeah, it’s true that what arm and x86 do here isn’t the best thing ever, just better than risc-v.
But to a first approximation, if you double the density of conditional branches in the program, then you will need to roughly double the size of the branch prediction tables to get the same performance, even if all of them are correctly predicted 100% of the time.
Currently only the most basic extensions are available. But nothing prevents RISC-V from introducing in the future an extension that extends the conditional code or an extension for integer/float overflow.
This point I didn't quite understand:
>Highly unconstrained extensibility. While this is a goal of RISC-V, it is also a recipe for a fragmented, incompatible ecosystem and will have to be managed with extreme care.
Most successful ISAs (including ARM) have their share of extensions, coprocessors, optional opcodes etc... ARM has the various Thumb encodings, Jazelle, VFP, NEON and more. Toolchains and embedded developers are used to dealing with optional features of computers, I'm not sure why RISC-V would fare worse here.
Beyond that I notice that many of the ascribed weaknesses are shared with other RISC ISAs like MIPS (but not ARM):
- No condition codes
- Less powerful, simpler instructions that require more opcodes to do the same thing but can potentially run faster.
- No MOV instruction
- The "unconstrained extensibility" is arguably a thing on MIPS too, with the four coprocessors that can be used to implement all sorts of custom logic.
Of course ARM has been more successful than MIPS, so maybe it's a sign that those things are indeed bad idea but given that this comes from an ARM dev I wonder if part of it is not just "that's now how ARM does it".
On the other hand I must say that I was surprised that RISC-V made multiplication optional, in this day and age it seems like such a useful instructions that it's well worth the die area. Optional DIV I can understand, but an ISA without MUL? That's rough, even for small microcontroller-type scenarios.
Having done just a tiny bit of compiler development for ARM, I can assure you that having all of these variants is a pain. Making compiler writers' lives harder means you're less likely to get optimal performance. At least on the more exotic variants, but possibly even on the most common ones.
I don't think this was meant as assumption about the authors gender. The same way i wouldn't assume that there is physical, actual pain involved when you said "having all of these variants is a pain" even when you literally wrote it.
Isn't it SVE? So far, it has been only implemented AFAIK by the A64FX, which is used by the current #1 supercomputer in the TOP500 list, but it wouldn't surprise me if we start seeing it on newer 64-bit ARM chips.
Note that SVE is a superset of Neon and SVE2 is a superset of SVE.
RISC-V's 2.2 (final, stable) ISA came out in 2017 and its already fragmented.
You do have different chips with different instructions but in a very regularized way. An ARM v8.2-A chip is going to have the same instructions whether it's made by ARM, Samsung, or Apple. And when v9 comes out they'll have NEON's SVE replacement everywhere and you'll be able to use the same code regardless of whether the SIMD width is 128 bits or 512.
I have yet to see any concrete details on ARM v9, generally speaking ARM v8 is pretty damn well designed I am wondering what v9 will look like.
https://community.arm.com/developer/ip-products/processors/b...
That's fragmentation right there. Severe fragmentation. Two completely incompatible ISAs.
Small 32 bit RISC-V comes in smaller and lower power than An M0, and small 64 bit RISC-V is not much bigger than an M0 and is rather popular controlling something in the corner of a larger 64 bit SoC.
I can empathize, but isn't that just part of the job of making a compiler? Any successful, long-lived ISA is going to have extensions and revisions that will need to be handled in the toolchain. I guess my point is not so much that it isn't painful, it's more that I don't really see what makes RISC-V really different besides the fact that it's a younger ISA and therefore we don't already know for sure which extensions are going to become de-facto standard and which ones will be less common.
>I believe the author doesn't identify as a "guy".
Arg, of course the one time I don't use gender-neutral language I manage to mess it up. Edited, thanks.
One of the best lessons I got when I was being inducted into the compiler club was: compilers are hard. It’s a hard job so other people can have easier jobs. It’s ok if compilers turn complex and managing that complexity is just something you have to learn to do. I don’t think it’s true that the need for that complexity leads to lower perf.
But I kind of assumed most ops perform well enough and important optimizations have a test case for and will get done
There are advantages to the RISC-V approach it is likely to lead to more fragmentation - and worse gives the ability for a major implementation to add proprietary extensions that are not licensed to anyone else putting smaller players at a disadvantage and leading to fragmentation, not only in the hardware but also in the software ecosystems.
Whilst you may not like ARM having control at least everyone (for a fee) has full access to the ISA and implementations.
I should add that I think that it's possible that Nvidia/ARM combination will remove ARMs 'level playing field' and we might see Nvidia only extensions for their designs - which would not be good. We'll have to see.
The same applies here: the more extensions and things you pile on... the less likely they are to get used unless they are de-facto mandatory. Even today you can see games getting released that won't touch AVX instructions on both Intel and AMD because of compatibility reasons.
Valve provides data to developers on penetration of various ISA extensions via the hardware survey ( https://store.steampowered.com/hwsurvey/Steam-Hardware-Softw... ) but for RISC-V there is no way to do the same. So most utility writers will be highly constrained in what they will use in terms of expected extension use, that will have significant harms in terms of performance. Alternatively it requires recompiling for every single different target, which is also likely.
In essence: you need a good baseline of compatibility for people to expect to use. It makes moving software easier between targets. A piece of software might certify for example on R64GC but not on R32IF because the double precision emulation might not work as expected or the lack of carry could be an issue etc.
I still wonder about RISC-V. To me, it seems pointless. But a lot of companies are buying into it so I'm wrong
Why would you ever want a standard ISA? If you're buying chips you either want a cheap standard one or a powerful efficient one. To be efficient (or cheap) you'd want to only support what's required and what works best with the implementation.
I don't really understand the point of a generic ISA. Why not have some kind of bytecode or standard format (like llvm-ir) that gets optimized for the CPU and gets a native binary that doesn't need interpretation.
Like how the f* is it easier to make something regular+generic fast rather than something custom for your hardware/chip/cpu fast?
Do you want to know how many times I used XML when it's not required? 0. Do you know how many times I used SQLite or my own binary file? I lost count. SQLite has far more constraints than XML and custom binary files/formats aren't hard after you done than a few times.
So are you here right now declaring that RISC-V and all those companies are in the wrong and risc-v will be a disaster?
Cause I might agree and be with you on that lol
-Edit- I have no idea what the state of the compiler is
https://github.com/riscv/riscv-gnu-toolchain
> Warning: git clone takes around 6.65 GB of disk and download size
WTF?
If you're going that complex than... wtf?
My reply was mostly aimed at the idea that you can move the complexity from hardware into compilers: it might be possible, but we know how to build out-of-order CPU better than we know how to build smart compilers, so you have to invest a lot more research time, and it's generally a lower priority.
Even innovations from the past decade or two, like VSDG, haven't made their way into "industrial" compilers yet.
As for the size thing:
GCC and Clang are huge (at least when you include their entire change history, which git does), RISC-V is comparatively only a tiny part of them, you should probably look into that further before jumping to conclusions.
You don't even need a whole separate toolchain with Clang or Rust, the whole "need to build GCC yourself to cross-compile" is outdated GNU tradition, not some kind of technical necessity.
1) Fusion is hard. While it can be hard (fusing x86 will be), I do not see why it would be meaningfully difficult for RISC-V hardware (except on devices so small it's better to have the simpler base ISA anyway), or compilers, who can in the worst case just treat fused pairs as their own instructions.
2) There's anything wrong with just most software assuming a fairly fixed set of extensions, as seems to have happened. If microcontrollers want to use a subset without multipliers, that doesn't mean anyone else has to care. If bitmanip is stabilized before RISC-V breaks into more common consumer use, why not assume it when writing code? It's only a problem if people make it one.
Most of the rest don't matter much in a global sense. The arguments about which operations go in which extensions might have meaningful merit, but it seems not very important to me.
This is why Intel publishes software optimization guides that go over what their fusions are, but it doesn't seem like the RISV-C spec is going to do that yet for many cases. And compiler authors need to know which instruction stream to generate to ensure fused execution.
There is always room for creativity, but that would be the same with or without indexed loads in the base instruction set. Any non-monopolistic hardware ecosystem has this problem; we've been able to ignore it largely on x86 since Intel had had a performance monopoly for so long, but once you have multiple competing core implementations compilers will have to worry about the edge-case performance differences.
What I'm talking about is more specific to groups of instructions that are safe to treat as fused by default. Note that even if the compiler outputs a pair of instructions but the hardware running the code doesn't fuse it, out-of-order execution means the penalty will generally be extremely small versus the best unfused instruction schedule.
RISC-V does give guidelines on which instructions are good fusion candidates. See for example section 2.13 in the bitmanip extension document.
Hardware, naturally, just has a fixed set of fusions it does.
- add or sub or mul
- 32 bit or 64 bit
- signed or unsigned.
I don’t remember which adventure you need to pick to get 5.
Source: I had to make most of these fast to make JSC competitive.
Edit: I said all, should have said most. Unsigned is less important for JS.
The ugly parts are indeed all ugly, though they have now added hint instructions.
ISAs without conditional moves tend to have predicated instructions which are functionally the same thing. I'm not actually aware of any traditionally RISC architectures that have neither conditional moves or predicated instructions. While ARMv7 removed predicated instructions as a general feature ARMv8 gained a few "conditional data processing" instructions (e.g. CSEL is basically cmov), so clearly at least ARM thinks there's a benefit even with modern branch predictors.
Conditional instructions are really, really handy when you need them. It's an escape hatch for when you have an unbiased branch and need to turn control flow into data flow.
The quantifiability comes from measuring results when you give compilers new instructions, vs paying implementation complexity (time, money and future baggage to support the insn forever). The upsides and downsides here come in different units so it's still tricky.
Lots of instructions can be proposed with impressive qualitative speeches convincing you how dandy they are, but in the end it's down to the real world speedup yield vs the price you pay in complexity and resulting second order effects.
(In rarer cases the instructions might be added not for performance reasons but to ease complexity and cost, that's where qualitative arguments still have a place when arguing for adding instructions).
It's fine if we don't have the evidence in this thread - I was just asking on the off chance that someone can point to a reference.
Seems the following have conditional moves: MIPS since IV, Alpha, x86 since PPRo, SPARC since SPARCv9
The following seem to omit conditional moves: AVR, PowerPC, Hitachi SH, MIPS I-III, x86 up to Pentium, SPARC up to SPARCv8, ARM, PA-RISC (?)
PA-RISC, PowerPC, ARM at least do a lot of predication and make a high investment to conditional operations (by way of dedicating a lot of bits in insn layout to it), but also end up using it a lot more often than conditional move tends to be used.
ISAs that evolved conditional move:
- MIPS
- SPARC
- x86
- POWER (isel)
ISAs that started life with it: - ARM (via general predication)
- Alpha
- IA64 (via general predication)The take away there is that general predication was found to be overly complex where the vast (vast!) majority of the benefit can be modelled with conditional select.
y = cond ? 0 : -1;
y = cond ? x : -x;
x = cond ? 0 : x+1; //< look ma, circular addressing!Observation re list of ISAs that evolved conditional move vs ISAs that omit conditional move: MIPS, POWER, x86, SPARC all targeted high power "fat core" applications at the point where it got added. AVR, Hitachi SH, PowerPC didn't add it while being driven more by low power / embedded applications. And many ISAs continued to see wide use in the pre-cmov versions of the ISA in embedded space (eg MIPS) after the additions. (PowerPC even removed it when being modeled after POWER)
This is not a direct answer to your question, but: I recently had to tune the conditional move generation heuristics in the GraalVM Enterprise Edition compiler. My experience has been that you can absolutely get decent speedups of 10-20% or more with a few well-placed conditional moves. The cases where this matters are rare, but they do occur in some real-world software, where sticking a conditional move in some very hot place will have such an impact on the entire application. Conversely, you can get slowdowns of the same magnitude with badly placed conditional moves.
It's a difficult trade-off, since most branches are fairly predictable, and good branch prediction and speculative execution can very often beat a conditional move.
JSC goes to great efforts to select it in certain cases where it’s a statistically significant overall speed up. I think the place where it’s the most effective for us is converting hole-or-undefined to undefined on array load. Even on x86 where cmov is hella weird (two operands, no immediates) it ends up being a big win.
Didn't check, but I suspect that decodes at least into two microinstructions.
Basically:
- LL/SC can prevent ABA if the ABA-prone part is in-between a LL and SC instruction
- To have a ABA prone problem you need some state implicitly dependent on the atomic state but not encoded in it. Normally (always?) the atomic state is a pointer and we depend on some state behind the pointer not changing in a context of a ABA situation (roughly ~ switch out ptr, change ptr target, switch back in ptr, through often more complex to prevent race conditions).
This means in all situations I'm aware of LL/SC only prevents the ABA problem if you at least can do one atomic (relaxed ordering) load "somehow" depending on the LL load. (LL load pointer, offset or similar).
But the RISC-V spec doesn't only not guarantee forward process in this cases (which I guess is fine) but goes as far as explicitly stating that guaranteed not having forward provess is ok, e.g. doing any load between the load reserved and store conditional is allowed to make the store conditional fail>
Doesn't that mean that if you target RISC-V you will not benefit from LL/SC based ABA workaround and instead it's just a slightly more flexible and potential faster compare exchange which can spuriously fail?
The spec says you are supposed to detect if it work and potentially switch implementations. But how can you do that reasonable if it means that you have to switch to fundamentally different data structures, which isn't something easily and reasonably done at runtime.
Or do I miss something fundamental?
> A path consists of a sequence of path segments separated by a slash ("/") character
Of course servers are perfectly free to do whatever they want, there is no HTTP police to stop them.
As long as Wikipedia doesn't use relative paths it is not a problem to have slashes in the URL.
https://en.wikipedia.org/w/index.php?title=Load-link%2Fstore-conditional
You can actually visit that if you like.But this is where the problem starts. E.g. for RISC when using LR/SC in a way which prevents ABA your are always losing all forward guarantees and it's totally valid for a implementation to be done in a way which will just never complete in such cases...
Botching the code can be done with either mechanism. Don't do that.
What gives me some hope is that an open hardware and Free Software world would help so many people, businesses, and governments. I think it would be a rising tide that lifted all boats except for specific tech industries.
That said, good article.
EDIT: Also, when the B extension becomes standard that should fix some of the issues.
JALR for call and return used the same opcode but are two different instructions, there is no need for "extra logic" for the decode or for branch prediction.
The lack of "register + shifted" could easily be circumvented by adding an extension for "complex arithmetic instructions".
And macro op fusion is a common solution that already exists in modern CPUs to increase the number of stages in the pipeline.
> Multiply and divide are part of the same extension
An extension can easily be partially supported in hardware (e.g. multiplication) and leave the other instructions emulated in software (e.g. divisions)
> No atomic instructions in the base ISA. Multi-core microcontrollers are increasingly common
But some microcontrollers do not need atomics. And if you are designing a microcontroller that does, just include the atomic extension.
Many of the criticisms made here are incorrect and are due more to a misunderstanding of RISC-V than to RISC-V design flaws
Any educated opinions how bad the RISC-V problems are compared to others when looking at the big picture?
Truly innovative designs don’t actually use RISC encoding and hide the details of instruction encoding. See for example NVIDIAs PTX , which gets mapped to hardware specific instruction streams, which are nothing like a traditional RISC architecture.
To me RISC-V is another example of how open source often produces lowest common denominator copy cat versions of ideas that are 30+ years old. Just like Linux. The sad thing is that this often kills innovation and locks in suboptimal designs for a long time, because it is hard to compete against something that is free.
Aren't patents limited to 20 years? Can't you just ignore all inventions that are < 20 years old to be safe from patents? Then you'd still be 10 years ahead of POWER.
The need to be easily reachable puts constrains on the design but is also a benefit in that many people will know this ISA and be able to make tools for it.
I think I would probably recycle the Alpha ISA circa 21164 (EV-5) with maybe a CAS instruction. It was pretty balanced between hardware and software and a lot of the complications in the VLSI design (dynamic logic, mostly) are moot with a modern technology if you stick with reasonable speeds.
Presumably now that the MIPS unaligned byte access patents are expired, a whole bunch of the idiocy that Alpha had to abide to avoid that patent can just be sidestepped.
One of the things that comes up with Risc-V a lot is code density, and it is a sore spot for many CPU designers because it is shamelessly abused by marketing departments as a measure of 'goodness'. This has sort of trained these engineers to flinch when something doesn't exhibit good code density, and the author is no exception.
However, FLASH/RAM volume has gone up hugely. This is in part because once you run out of logic to lay down in a chip you flood fill the rest with RAM and/or FLASH because hey the chip has to be big enough to hold pad landings for all of its pins. This has taken a lot of pressure off the code density thing and now it seems appropriate to look at algorithmic capacity.
I agree with the author that it is a much better use of resources to put a hard multiplier on a chip than it is to have more space in flash so that you can do that with instructions. But from a algorithmic capacity question? It is all Turing complete so really what is the practical difference?
Now I started life programming on PDP-11's that had a "native" instruction set that was not unlike RISC-V in being maximally simple. We called it "microcode" :-) And the "real" instructions were actually sequences of microcode in a microcode ROM. That was pretty cool because you could swap out floating point instructions for vector instructions or string handling instructions if you wanted.
I will not be surprised in the least if people design "custom instruction sets" that layer on top of a RISC-V core, just like layering a front end stack on web assembly. And the available resources to do that are pretty plentiful.
One of the coolest architectures I got to play with was the Xerox "D" machines. They took this to an extreme and some amazing software was written for them. You could get really close to maximizing utilization. It was very interesting to load a new instruction set, recompile your Mesa code, and have it run faster with no changes to the hardware at all.
Of course DEC and Xerox didn't invent this, the IBM 360 had a big button on the front "IMPL" which was "Initial Micro Program Load" which prepped the instruction set for what ever OS you were about to start up.
And now you can build systems like this with open source tools and off the shelf FPGA dev boards. Such a great time to be interested in systems architecture.
[1] https://twitter.com/erincandescent/status/115453579942322995...
I don’t agree. I’m working on a chip right now where NVM and RAM is extremely constrained. There will probably be many millions made, so it’s still relevant.
On the last chip I worked on, it was perhaps true to some degree. But that chip, even though it was a microcontroller, had a decent cache. Access to NVM is slow. And code density can have a significant impact on cache performance.
Quite often in some applications, for example hash tables (a very popular data structure).
:-)
(No, the M extension does not help.)
I mean, whatever the exact goal of RISC-V, I don't see any reason why an analysis and critique would not be relevant. Nobody invests resources into this with the intention that it gets ignored.
Also, Wikipedia has a list of RISC-V implementations.
For each instruction, I would guess that a 'draft' compiler is produced that can use that instruction in its code generation steps, and a few 'draft' CPU designs are made which include support for that instruction.
Then cycle accurate simulations can be done on a set of test benches to see how the addition of that instruction affects performance, power, and code size across a wide array of different usecases.
In the case of RISC-V, where some CPU extensions might be emulated, I would expect the test to also cover the performance hit of emulation of the extension for those machines without native support.
If all of that was done, and still it made sense to add the instruction, then most of these critiques aren't valid - since there will be hard data that the approach taken was the best one.
Perhaps when RISCV was a young project, too many design decisions were made without the massive compute farm to do all these simulations, or before more complex multi-issue CPU designs were added to it, and therefore some decisions aren't optimal?
Small ISA != Small transistor count.
People will inevitably try to throw more transistors on the ISA limitations.
They are all kinds of changes in flow, it's not crazy to have an instruction for that primitive.
He complains Risc-V need 4 instructions to do what x86_64 and arm does in two, but... it says Risc-V. And x86_64 CISC instructions devolve to a pile of microcode anyway.
Guy invested in powerful incumbant using a completely different and even more established enumbant to bash the challenger over the head with doesn't feel like I am learning anything useful when I read it.
(puts on chip designer's hat) Essentially the ARM/x64 case turns the load memory address calculation into a 4-input adder (so 2 layers of adders) and maybe a couple of extra gates because there are multiple addressing modes. Riscv's equivalent is a 2-input adder. Those get into a critical cache (and TLB) access path and that limits how fast your CPU's core clock can be (or forces you to split that path into 2 clocks).
Essentially that's part of the whole RISC idea - simple means faster - you can run your core clocks faster if the decode (and address calculations etc etc) are simpler - getting rid of lots of addressing modes was a big part of the original RISC movement. I think all 4 of those riscv instructions are 16-bit ones so they may even fit into the same space as the 2 ARM ones (haven't hacked on ARM for a while)
BTW chances are that that x86_64 mov instruction is not being devolved into more than one internal uOp (might be two if they separate off the address calculation into it's own uOp)
Nobody believes this anymore, not even the RISC guys. Look at Apple chips; the advantage of a simpler instruction set is width (degree and depth of superscalar execution) and the ability to put more optimizations in hardware. Clocks are all bound by roughly the same limits nowadays.
The problem with this is that today, your clock speed is bound by neither your decode nor your ALUs. Only implementing weaker ALUs makes sense from an optimization standpoint if it buys you more clock speed. But as it doesn't, it just leaves you competing with another CPU that has the same clock speed as you do, and which does a lot more per clock than you do.
But at other pointed out this only makes sense if the ALU is the critical path and that having a two level adder significantly impacts the latency
How does this philosophy fare now that we're generally at the top end of feasible clock speeds for processors, given realistic power and related cooling budgets. Intel and AMD (and I assume others, but I don't follow them as closely) have been increasing throughput by doing more per clock.
It would seem VLIW is the recipie for simple hardware, doing more per clock cycle; but that hasn't had good results either.
The only real benefit is that fixed length instructions are easier to decode. That's about it. If you can get away with fewer instructions then that's what you should do.
So… what, it should take 5 instructions?
Executing more instructions for a (really) common operation doesn't mean an ISA is somehow better designed or "more RISC", it means it executes more instructions.
>And x86_64 CISC instructions devolve to a pile of microcode anyway.
Some people seem to have this impression that like every x86 instruction is implemented in microcode (very, very few of them are) and even charitably interpreting that as "decodes to multiple uops" (which is completely different) is still not right. The mov in the example is 1 uop.
True. But as bonzini points out (or rather, hints at) in https://news.ycombinator.com/item?id=24958644, the really common operation for array indexing is inside a counted loop, and there the compiler will optimize the address computation and not shift-and-add on every iteration.
See https://gcc.godbolt.org/z/x5Mr66 for an example:
for (int i = 0; i < n; i++) {
sum += p[i];
}
compiles to a four-instruction loop on x86-64 (if you convince GCC not to unroll the loop): .L3:
addsd xmm0, QWORD PTR [rdi]
add rdi, 8
cmp rax, rdi
jne .L3
and also to a four-instruction loop on RISC-V: .L3:
fld fa5,0(a0)
addi a0,a0,8
fadd.d fa0,fa0,fa5
bne a5,a0,.L3
This isn't a complete refutation of the author's point, but it does mitigate the impact somewhat.But hash tables are used here and there, also in loops.
Some people know them as "dictionaries" or "key/value stores".