So far RISC-V has failed miserably at both. "Just solve it in decoder", "just fuse common idioms", "just make vsetvli fast". Sure. Not like designers of ARMv8 or X86S had decades of prior experience to make better decisions.
So far RISC-V has failed miserably at both. "Just solve it in decoder", "just fuse common idioms", "just make vsetvli fast". Sure. Not like designers of ARMv8 or X86S had decades of prior experience to make better decisions.
It’s far easier to add things to interfaces than remove them.
One mistake they have made that may bite them later is their compressed instructions extension. It takes up far too much of the instruction set space for the amount of utility it provides, even on microcontrollers in my opinion. It also introduces lots of edge cases around word alignment across cache and memory protection boundaries.
The RISC-V C extension takes up 75% of the opcode space.
The 16 bit Thumb1 instructions take up 87.5% of the ARMv7 opcode space.
Not to mention predication taking up 92.75% of the classic ARMv1-ARMv6 instruction set.
RISC-V has comparatively twice as much "32 bit" opcode space as ARMv7, and 3.45x more than classic ARM A32.
Here is the RISC-V design as at 13 May 2011:
https://people.eecs.berkeley.edu/~krste/papers/EECS-2011-62....
Compare that to the design eventually ratified in 2019 and you'll find some changes in details, but not in concept.
- instructions encodings changed to e.g. put rd on the right next to the opcode and the MSBs of literals on the left.
- J and JAL are now combined
- JALR had three versions (in func3) to distinguish call / return / others. This is now done as a convention which registers are used in rd and rs1. (It matters only for a return address prediction stack, an advanced feature)
- RDNPC (get address of next instruction) later turned into AUIPC (which includes adding an offset to the PC)
- on the other hand, it already has the very RISC-V feature of loads using rd and rs1 while stores use rs1 and rs2 (and the offset split up differently). Most other ISAs (including all Arm ISAs) use the same register fields for loads and stores and the offset in the same place, resulting in either loads using rs2 as the destination or stores using rd as a source!
Seriously, other than the changes mentioned above, this document looks just like the final RV32G/RV64G spec from ratification in 2019. (I didn't check the floating point instructions in detail).
As for SIMD -- RISC-V was designed as the control processor for an advanced length agnostic vector processor. F*ck fixed length SIMD.
There is the issue. It goes completely against how the current landscape of SIMD support looks like and what everything is optimized around (it's fixed length SIMD, variable length has huge issues getting mapped onto anything that isn't auto-vectorization).
Now, I'm not a fan of RISC-V so, if anything, I encourage their way of doing RVV. Makes it more likely to die.
Much as I like RVV, fixed-width use-cases still exist and are pretty important for CPU SIMD; scalable vectors work well for very-many-iteration embarrassingly parallel loops, but those are also things most suited for being moved to a GPU. Where CPUs have the most potential is in things with some dependency chain or small loops, for which scalable vectors largely just add questionability of performance.
The developers of the dav1d software AV1 decoder found that RVV on a Kendryte K230 [1] performed no worse than NEON on an A53 on their small fixed-size transforms on video CODEC blocks.
https://www.youtube.com/watch?v=asRnBcn5VKs&t=9m40s
[1] THead C908 core, very similar to the old C906 core found in e.g. the $3 Milk-V Duo, but dual-issue and with the RVV updated from draft 0.7 to ratified 1.0
It very well could; for example, RVV hardware with VLEN=512, one 512-bit vector ALU, and three scalar ALUs (float or integer), would have 512 bits/cycle with vector and 192 bits/cycle for scalar (i.e. vector at VLMAX is beneficial!), but using only 128 bits of the vectors would end up with just 128b/cycle. And this is just regular ALU instructions, vrgather can get significantly worse, and it is extremely important for most fixed-width stuff.
This of course wouldn't ever happen with actually-128-bit hardware as it wouldn't ever make sense to have such vector vs scalar distribution for its guaranteed-128-bit workloads. And, though perhaps my previous example is a tad extreme, with higher VLEN it gets more and more reasonable to have a larger gap between scalar and vector ALU/port counts, and maintaining 'vector_alu_count*2 ≥ scalar_alu_count' gets less and less reasonable.
K230 is irrelevant here - it has VLEN=128, which is the optimal thing for being used as a fixed-width system / imitating NEON. And, more generally, looking at just one RVV implementation cannot give any insight on how well the "scalable" part of RVV works, as that only applies when running the same code across multiple different implementations with different VLEN.
Could high-VLEN hardware still be made such that 128-bit usage is never strictly worse than scalar? Perhaps. But there exist incentives under which it might not make sense to have such, which have no possible equivalent on NEON/x86, and apply specifically when the "scalable" part of RVV is actually used for its intended purpose in hardware.
As for the longer vectors, no way to say for sure until we have such hardware, and even then different vendors will have different goals and very likely different quality of implementation.
At the time they ported that code the K230 was the only RVV 1.0 chip available.
The SpacemiT K1/M1 with 256 bit vector registers has now been available in the Banana Pi BPI-F3 for a few weeks, and is about to ship (this month they say) in the Sipeed Lichee Pi 3A, Milk-V Jupiter, SpacemiT Muse Pi. So no doubt the dav1d people (and others) will be getting busy with that.
Around the end of the year we'll have SG2380 machines with SiFive X280 cores with 512 bit vector registers. That should be very interesting, as SiFive have been designing those since at least 2018 and the quality of implementation should be very good.
> and even then different vendors will have different goals and very likely different quality of implementation
And that's, like, the entire problem. RVV's design allows for otherwise-reasonable designs where fixed-width usages suffer (whereas there's literally no reason to do equivalent things on NEON/x86). And given that I doubt that the designers of such chips would be making legal mandates to not run general-purpose software on such, were such hardware to become real, general-purpose software could easily end up being expected to work well on it.
Yes, said RVV design does in fact allow having those different goals, which is good. ..For those goals. Pretending noone's gonna try to use hardware for anything other than its strictly intended purpose is nothing more than wishful thinking.
SG2380 / X280 are already known to have a very horrible vrgather, which is gonna be extremely awful for fixed-width stuff, and even some scalable stuff too. And SG2380 is rather explicitly for general-purpose computing.
Which is, basically, what Intel or Arm do, at any given time.
I prefer the possibility that some people make amazing products, and some people make underperforming (but possibly much chaper) products, and the market decides which ones to buy. The most important thing is that they can all use the same roads (software), even if they should stick to different lanes.
> SG2380 / X280 are already known to have a very horrible vrgather, which is gonna be extremely awful for fixed-width stuff, and even some scalable stuff too.
I don't believe that to be the case. IIRC X280 does vrgather in ~1 cycle for vector lengths up to the datapath width (256 bits, 32 bytes), so anything corresponding to your NEON or AVX2 cases are going to be just fine. As I'm sure you know the computational complexity of vrgather is proportional to vl^2 so no one is going to do single-cycle vrgather at LMUL=8 => 4096 bit sizes, or even probably at 512 bit.
You could use a 256 bit datapath to implement 512 bit vrgather in 4 cycles, and LMUL=8 (4096 bit) in 256 cycles but that's not the intended use for the X280 so they didn't spend those transistors.
At longer sizes X280 does 1 element per cycle, which is still a significant speedup over what you could do on the dual-issue in-order scalar CPU.
> And SG2380 is rather explicitly for general-purpose computing.
The *P670* cores on the SG2380 is for general-purpose computing. The X280 cores are for media processing and similar tasks that don't typically use large gather operations.
> As I'm sure you know the computational complexity of vrgather is proportional to vl^2
But that sucks for anything wanting just a fixed 16-byte table (which is a pretty frequent thing in fixed-width SIMD), and low VL will not help with that as the table argument of vrgather is always VLMAX; you necessarily need low LMUL (so a more correct average complexity is LMUL*VL even for hardware dynamically scaling resource usage based on VL). But, if I wanted a 32-byte table, I'd have to jump to LMUL=2 for portable code, at which point VLEN=512 hardware would be forced to consider each result element potentially selecting from 128 bytes of data.
Granted, here at least it can be reasonably expected that general-purpose hardware would make LMUL=1 not horribly slow for typical use-cases (be it via spamming silicon at it, not having high VLEN, or specializing for small local variance of indices at runtime (which, while neat, is silicon that could've been spent on actually meaningful things instead of reconstructing what the developer already knew; and even then it probably would result in higher latency)).
Unlike with choosing between SSE, AVX, AVX512 etc, with RVV you don't have to duplicate your code to do so.
Whereas with RVV you might have to dynamically select LMUL to not get unreasonable perf (a very low bar!).
And.. a 16- or 32-byte shuffle should not be considered as some "critically tuned" thing, this is basic fixed-width SIMD stuff.
And if hardware making bad tradeoffs for fixed-width usages ever makes it to general-purpose usage to any significance, RVV as a whole would end up being required to be considered as not fit for fixed-width stuff (outside of extremely important & well-funded things where people can afford to add manual tuning for each separate funky CPU, but it's a pretty safe bet that most developers won't have every single piece of RISC-V hardware.. especially that which they would hate).
Without reasonable performance guarantees, "RVV can imitate NEON" is as useful of a statement as "base RV64G without any extensions at all can imitate NEON". It's just stating that it's a turing-complete system.
And, again, even at 1 element per cycle, using vrgather is significantly faster than not using it and using the dual-issue scalar side instead.
AND at short NEON or AVX vector lengths the X280 is single-cycle for vrgather *anyway*.
But, sure, X280 might not be intended for general-purpose usage. Still pretty sure that won't stop people from trying. And it's far from impossible, and, to some extent even reasonable, for future hardware to be intended for general-purpose usage while still disadvantaging fixed-width.
And, for a good number of fixed-width SIMD usages, the alternative might not necessarily be "do everything exactly as with SIMD but in scalar registers with one instruction per would-be-SIMD-element"; namely, SWAR could be utilized, some shifts/rotates/bswap/multiplies used in place of shuffles, entirely different algorithms used, and you wouldn't need to duplicate computation for something like a cumulative sum/xor that on SIMD would've been implemented as log2(n) slides. I've had a good number of cases where some fixed-width x86 SIMD thing is only marginally faster than the scalar baseline.
I'm not saying this is some utter massive disaster that'll kill RVV; but vrgather definitely, as-is, is pretty inadequate for many things that it can actually do purely due to forcibly having both operands have the same LMUL and having only that as the cap on the table size; and it is unquestionably possible (though not guaranteed, and maybe even unlikely) that in the future software will be expected to handle funky-tradeoff hardware, which'd suck for software developers.
The proper solution would, I think, be to have duplicate `vtype` and `vl` CSRs that are set up by a different `vsetvl2` instruction. It would be seldom-enough used to dispense with the opcode-expensive immediate form and only take the type from a register, making it a cheap I-type instruction.
This would also solve the problem of table and indexes having different element sizes that prompted the late creation of the `vrgather16` instruction and imposition of the artificial 65536 bit limit on `VLEN` in RVV 1.0 and foreshadowing of a potential future `vrgather32`.
While vrgatherei32 is certainly definable, it not existing does have the benefit of not burdening small impls with yet more connections for indices. Also, vrgatherei64, spooky.
I do agree however, this would result in bad performance for VLEN>128 and VLEN<DLEN implementations.
Adding a gather instructions limited to 16 elements might be a solution, however it could also result in vendors feeling more free to have a slower LMUL=1 vrgahter, that still needs to perform reasonably. The LMUL=1 and 16 element vrgahters are the most common cases.
IMO the software ecosystem needs to steer the hardware here, in place of ARM for neon, and Intel/AMD for AVX, stopping people from implementations with disproportionately slow permutes. People can't expect to put a ML core, like the X280, as the main CPU on a regular desktop class processor and expect general purpose software to be optimized on it.
I suppose this could be added to the RVA profile as well, but they don't want to specify micro architectural details, so what does it mean for an instruction to be fast? Maybe there could be a non-normative clause that software is expected to use LMUL=1 vrgather for LUTs and other shuffles inside of hot loops. I suppose creating an issue regarding this can't hurt, so I'll look into it.
If that is true, then their higher LMUL implementations are basically a design mistake. They could've used the LMUL=1/2 vrgather to implement LMUL=1 and above by calling it repeatedly (LMUL*2)^2 times, that should add almost no extra die area. This is what the C908 and X60 seem to do.
I've written what I expect sane implementations to do here: https://gitlab.com/riseproject/riscv-optimization-guide/-/is...
This didn't take into account higher SEW, but I could imagine an implementation that as a LMUL2 native implementation for e32/e64.
The number of connections for a specific SEW and VLEN should be (VLEN/SEW)(SEW/8), however the regularity and distance probably also impacts final implementation cost.
This would result in k*VL/DLEN uops for simple index patters such as zip/unzip/reverse/broadcast with k=1 or k=2, and worst-case (e.g. transpose) is as bad as your option 2 with N=1 or N=2, and requires just DLEN^2 connections in the main shuffle, as much of the heavy lifting is done by the already-necessary silicon for getting DLEN-sized chunks of operands (maybe there might be complications with having those not being requested sequentially though; I'm not a hardware person).
Maybe this is what you had in mind, but I must admit that I didn't fully understand how you described the implementation:
For LMUL>1 vrgather:
Create one uop for every LMUL=1 register:
* Look at first index (n=idx[0]/(VLEN/SEW)), to select vector register to read from
* do the LMUL=1 vrgather from that register, and check if all are in range
* if they weren't, emit LMUL=1 vrgather uops corresponding to the other LMUL=1 register sources
This would give you LMUL cycles for all permutations that read all values for a LMUL=1 register destination from a single LMUL=1 source register, or the proportional equivalent when DLEN<VLEN. So this would cover the LUT/zip/unzip/reverse/broadcast cases.
As I mentioned I don't think this is a good fit for ooo designs, but in-order ones, with larger vector lengths could/should probably implement this.
https://github.com/llvm/llvm-project/issues/79196 also mentions this in "we want the elements to come from as few DLENs of the source vector as possible"; and down the comments there is me having had too much fun writing code to fill in "I don't care" index elements such that they don't add any more required chunks, in a completely-DLEN-agnostic way.
I don't think this is necessarily that bad for ooo; it's really "just" variable latency, which would already exist to an extent on cores which dynamically scale based on VL (though, granted, not all would be such, and the exact latency would be found much later down the pipeline even then).
It's not worth implementing when VLEN>=DLEN, beacause then you can replace all of the uses were you'd always get an atvantage with unrolled LMUL=1 instructions. So I don't think it will be common in ooo designs, since they usually have smaller vectors.