Note macro op fusion is widely used for other architectures already, particularly ones like x86 where what the processor actually runs looks nothing like the machine code.
Note macro op fusion is widely used for other architectures already, particularly ones like x86 where what the processor actually runs looks nothing like the machine code.
It doesn’t matter if they’re fused or not if the reduced instruction density increases memory usage and puts more pressure on I$.
Also, I don’t buy the whole fusion argument on the grounds that having to fuse super complex (5 instruction or more) sequences adds enough complexity that you’ve got opportunity cost. Much better for everyone if the CPU doesn’t have to do that fusion. That’s the whole point of good ISA design - to prevent the need for fusing in cases you’re doing something super common.
Edit: Here's the thesis about the design decisions in the C encoding: https://people.eecs.berkeley.edu/~krste/papers/waterman-ms.p... See also the diagram on page 62 of this document: https://people.eecs.berkeley.edu/~krste/papers/EECS-2016-1.p...
The thing I see is this: you can also add compressed encodings for any ISA. And that has its own costs (it’s harder to decode and it’s harder on software that wants to do bidirectional analysis of machine code). So “my isa has shortcomings but it’s cool because compression” isn’t a perfect argument since if your isa lacks those shortcomings then you still benefit from compression and you don’t need it as much, which is better.
RISC-V compact instructions don't require special modes and run in fully-mixed mode with 32-bit instructions without all the penalties thumb has (they are literally just extended into their 32-bit counterparts internally).
High code density became less valuable, was the point.
(IIRC they removed that one on x86-64)
not necessarily, since
1. the compressed versions are basically 1-1 to the non-compressed versions
2. the uncompressed versions are more for study/academia and clarity; its expected that irl only compressed instructions are used (for instructions that can be compressed)
more info: https://riscv.org/wp-content/uploads/2015/11/riscv-compresse...
Can you give an example of someone advocating for 5 instruction fusion? Normally it's limited to three.
You need fast sequences for all of these variants:
- add or sub or mul
- signed or unsigned
- 32 bit or 64 bit
Some of those need 5 instructions. I don’t remember which adventure you need to pick to get 5.
sadd32, sadd64, uadd32, uadd64, ssub32, ssub64, usub32, usub64, smul32, smul64, umul32, umul64
I don't think it's a fusion problem; even if you did fuse these sequences, they'd still be bad, since they'd be writing lots of extra registers.
Bitfield insertion is only one instruction in most RISC ISAs, but 5 or more in RISC-V.
That said, until it gets ratified by the consortium and implemented in silicon its still just a (well-researched) wishlist.
add t0, t1, t2
slti t3, t2, 0
slt t4, t0, t1
bne t3, t4, overflow
So two extra registers needed also..edit: so I now think a good extension to add overflow checking to RISC-V is with an instruction that works like "slt"- call it "sov", set if add would overflow:
add t0, t1, t2
sov t3, t1, t2
bnez t3, overflow
add/sov could be fused..The only way to know there isn't a dependency is if that register gets clobbered by something else very soon afterwards.
But this whole topic of "checking for signed overflow is expensive" is overblown. It's simply not that important an operation, especially in the context of those languages that do it a lot also doing a lot of memory references, which are far more expensive.
Adding arbitrary completely unknown integers is pretty rare. If you know both numbers are greater than zero then a single compare-and-branch is all you need. If one of the numbers is a constant then a single compare-and-branch is all you need.
auipc lr, zero, .LONG_TARGET20
jalr lr, lr, .LONG_TARGET12
Only one register (the link register) is clobbered, so the pair can be fused into a single wide jump-and-link.So in parent's example, sov might fuse with the following bnez, but it likely wouldn't fuse with the preceding addition.
Quite a few things determine whether a fusion is doable or not. In addition to the number of destination registers, you do, to a more relaxed extent, care about source operands, but also things like ‘does this fit nicely in a single pass through the pipeline?’ and even just ‘is this materially beneficial?’
Lots of cores (but not all) can write two registers from a fused instruction, given the right conditions, and sov does rerun the addition, so add-sov fusion sounds very doable to me.
(Quoted from OPs comment)
This isn’t a subject I’m an expert at but wouldn’t this mean there’s already some sort of translation going on already on the other systems? So it’s mostly just added end user work not a giant performance loss on RISC?
It would just then simply be a layer of abstraction that is lost.
The x86 front end breaks down big instructions into smaller RISC-like micro-ops and then fuses/re-orders/optimizes/etc and runs those instead. There’s pros and cons, the con being sheer complexity and power budget, the pros are that it’s an abstraction so the microarchitecture can change completely without recompiling your code —- and you get CPU specific optimizations too. The CPU is basically emulating x86.
You could in theory build an x86 CPU with a RISC-V core behind that decoder.
This is a widely held meme, but the internet at large doesn't have any evidence to back it up. A couple of publicly visible engineers that do have experience are on record as saying that cell-phone-class competitive x86 was absolutely possible. Intel and AMD chose not to pursue those markets.
The expensive parts of a high-end CPU aren't normally in the instruction decode part. They are in the branch prediction, branch mispredict recovery, forwarding networks, memory re-ordering, and so on. Anything short of a dataflow ISA has little impact on those structures.
Weren't there Windows Phone devices with x86 SoC, but they weren't competitive?
Commercially made x86 Android phones exist, the most popular to my knowledge were some of the Asus ZenPhone models.
When Atom was released in 2008, the A9 had already been announced (a year before). A9 was around 10-15% faster per clock and were often multicore meaning a 1.5GHz chip was faster in all metrics over the 1.6GHz Atom.
A few articles came out June/July of this year with Analysts saying that Intel had spent over 10 Billion dollars trying to break into the mobile market with no success. ARM's current R&D budget (according to Nvidia a month or so ago) is 0.5B. If they spent that much every year since 2000, they would barely match Intel, but their entire market cap in stayed under 4B all the way until 2009. Remember, that R&D includes their high-end ARM cores, but also GPU, midrange designs, various microcontroller designs, a realtime OS, ARM tooling, NPUs, various kernel support, etc
If that much money can't fix up x86 to keep up with a budget a fraction of the size, I take that as proof that the ISA really does matter.
[1] https://www.usenix.org/system/files/conference/cooldc16/cool...
I don't think you can generalize from that result to much of anything.
That struck me as one of the few apples to apples comparisons ever of the instruction sets at the high end, from a party not really incentivized to bend the truth one way or another.
But it could have easily come from something like the relaxed memory model, or they could have just been overly optimistic. The chip was cut after all.
With the exception of Atom, all recent Intel designs have been a RISC core with a CISC decoder slapped on top. Everything else being equal, the simpler decoder will create a smaller chip. Because the decoder is always running all-out, the simpler decoder will also use less power.
x86 instructions are multi-length from 1 to 15 bytes. The cost to slice that up is always going to be bigger than fixed-length instructions. RISC-V has variable length in theory, but in practice, compact instructions extend into 32-bit instructions with some bits added which is important for decode cache while longer instructions are ignored.
Because there's a maximum tolerable decode latency of a very few cycles and latency increases with cache size, decode cache size has a definite cap. x86 has a couple orders of magnitude more potential instructions than RISC-V. More instructions translates into a lower hit rate for the same size cache barring any heuristics (more on that below).
Matching variable-length arrays to an unknown set of arrays in cache is inherently a hard problem. Every solution has tradeoffs and the resulting heuristics are bad for computing (see below). In contrast, searching for a match on a fixed 32-bit array has much more simple general solutions that don't require tradeoffs.
C and CPU designs feed off each other. Let's say there are instructions X and Y which can do equivalent things. x86 engineers played around with both and got a lucky insight into how to make X a bit faster. Compiler writers jump on it and start using the faster solution. x86 engineers now all but stop looking to improve Y and spend their time tinkering with X instead. Compilers now focus even more heavily around not just X, but any instructions more closely associated with X.
In that entire (true) story, nobody gave a second thought to whether the final result of Y would have been faster overall if not for the lucky break with X. If x86 developers were actually free to choose whichever instructions they wanted, x86 decode would be much, much slower than it appears to be. This self limitation argues that perhaps a more RISC-like ISA is inevitable.
A new ISA where everything is used would definitely have complexity, transistor, and power disadvantages vs a new ISA that didn't make that mistake.
Don't think its strictly accurate to say that Intel 'chose not to pursue' the mobile SoC market - IIRC they tried, made little progress and gave up having spent a lot of money in the process.
Can't wait to see more information about the goldmont microcode work to see if that holds for intel as well as it does for AMD.
// &(array[offset])
slli rd, rs1, {1,2,3}
add rd, rd, rs2
The sequence in the article uses what Intel calls the fast case but it still wouldn't qualify for Berkeley's two instruction fusion. Dunno if anyone does three instruction fusion.As an aside, LEA is never getting added to the base RISC-V nor should it be. But I'm surprised it isn't considered for an extension.
[1] https://www2.eecs.berkeley.edu/Pubs/TechRpts/2016/EECS-2016-...
I'm not sure that is the idea, given that RISC-V is targeted at processors so low-end that they don't even implement multiply.