Programming with RISC-V Vector Instructions
gms.tf
gms.tf
Having worked on PPC, with a RISC Instruction set the difference between 128 bit vs 256 bit instructions really does eat up the limited opcode space for really trivial differences.
Also having written say one fast version of vector array copy can now just be used between different vector version lengths, no need to write different versions to exploit expanded width, and same goes for many vector compiler optimizations.
How does this work with physical registers as opposed to architectural ones? Typically in PPC the 128 bit and 256 bit ones were architecturally and physically non-overlapping so you did get extra registers when you go from 128 to 256 or 512. I don't know if that's the case for RISC-V here.
But yeah, brilliant looking forward to more!
The grouping... sounds “interesting” to implement in an OoOE design. Most obvious would be to have the instruction decoder emit one uop per register in the grouping... but that means vsetvli would have to stall decoding until it’s resolved. But that also seems to be how element size is set, so that would kill the performance of mixing precision in the same kernel...
Well I guess it could assume the grouping doesn’t change and flush the pipeline if it did. But you still don’t want to be mixing kernels with different groupings...
It is expected that in a future expanded Vector instruction set with 48 bit or 64 bit opcodes the vtype will be explicitly encoded in every instruction and can change every instruction.
Right now the vsetvli is setting (slightly) persistent state that affects following instructions. You're allowed to put one before every vector instruction if you want, without significant execution penalty -- there will be a little, from extra instruction fetch and decode -- but similar to doing, say, an integer add between each vector instruction.
The natural implementation even now is to have each vector instruction pick up the current vtype when it is decoded and carry it along with it through the pipeline as a few extra bits of opcode.
You certainly don't want to have any stalls or pipeline flushes just because the vtype changes.
> For the purpose of our example, the exercise is to write vector code that efficiently converts a BCD string such as { 0x12, 0x34, ..., 0xcd, 0xef } to a corresponding ASCII string (e.g. { '1', '2', '3', '4', ..., 'c', 'd', 'e', 'f' }). On a high-level, a solution involves separating the nibbles into single bytes and then converting each byte to the matching ASCII value.
If your BCD string has 0xcd or 0xef in it, it's not BCD, is it? It's "binary coded hexadecimal", or as we usually call it, "binary".
This code converts a byte string to its hex representation. It has nothing to do with BCD, right?
BCD to ASCII is a strict subset of bin to hex ASCII; and in this case there is no runtime cost to supporting both. This also covers nybble-coded octal.
A couple of things need to be changed to bring it up to date:
-vlbu.v v16, (a1)
+vle8.v v16, (a1)
+vzext.vf2 v16, v16
-vsb.v v24, (a0)
+vse8.v v24, (a0)
I'd also probably make (or at least compare the speed of) one more change: -vrgather.vv v24, v8, v16
+vmsgtu.vi v0, v16, 9 # set mask-bit if >9
+vadd.vi v16, v16, '0' # add '0' to each element
+vadd.vi v16, v16, 'a'-0xA-'0', v0.t # masked add to correct A..F
That's basically the same code as he used to create the lookup table in v8 for the vrgather. I think it might run faster on many machines, and also the vrgather would fail on the smallest machines (with only 32 bits in each vector register) if anyone modified the code to not use m8 (LMUL=8).For more explanation see my post at https://www.reddit.com/r/RISCV/comments/i5alno/programming_w...
[NB slight cheat -- those vadd.vi immediates are too big to fit .. you'd actually need to put them in integer registers and use vadd.vx, which can be set up outside the loop]
EDIT: Non-layman wording
> Thread contexts with active vector state cannot be migrated during execution between harts that have any difference in VLEN or ELEN parameters.
https://github.com/riscv/riscv-v-spec/blob/master/v-spec.ado...
Additionally, I believe experimenting with vector ISAs was mentioned as one of the reasons they started another RISC research project, which ended up being RISC-V.
https://people.eecs.berkeley.edu/~krste/papers/EECS-2016-1.p...
BTW if unaligned 32 bit instructions are a concern there is a Compressed NOP (C.NOP == addi x0, x0, 0 but without RAW hazards).
It's simply Not That Hard to deal with. You just need to have two 32 bit words in your instruction decode buffer. Sometimes you need the 1st half of the 2nd word and sometimes you don't.
Incidentally, once you've done that, arbitrarily aligned (on halfwords) 48 bit instructions don't need anything extra.
vsetvli t0, a6, e8, m8 # switch to 8 bit element size, # i.e. 4 groups of 8 registers
vmsgtu.vi v0, v8, 9 # set mask-bit if greater than unsigned immediate # --> v0 = | 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0 |
If vsetvli results in groups of 8 registers, then surely vmsgtu.vi only affects v0 which is the first 8 registers? The following 8 are in v8 if I understood the previous writing correctly.
(What personally don’t understand is the point of these register groupings; they seem a bit extraneous and error prone, as you already can set the element size and you have to guess at the minimum vector register size…and why wouldn’t you always just set them to 8? I think ARM’s SVE does something similar but it fixes the size for you essentially.)
The current RISC-V "V" draft standard requires the vector registers to be at least 32 bits wide (VLEN ≥ SLEN ≥ 32)[1], so a 128-bit vector with 16 8-bit elements may require up to four registers. Setting the group size to eight is a bit extravagant, but since the total number of elements was limited to 16 any extra registers in the group will not be affected. With a smaller group size you could potentially use those registers for something else.
P.S. I think you have the mask bits reversed. The instruction is "set if greater than" so v1 should be 0x0101010101010000 and v0 should be 0x0000000000000000 (corresponding to the |1,1,1,1,1,1,0,0,0,0,0,0,0,0,0,0| state shown in the article for the combined v1:v0 register group).
If you're simulating actual hardware, e.g., a Verilog description of a RISC-V microprocessor compiled to a cycle-accurate simulation with Verilator, your simulation rate is going to be ~10KHz. You can write useful tests with the Proxy Kernel (or something like it) that run in ~1 million instructions (minutes of wall clock time) while still getting full system calls like printf. However, booting Linux is out of the question (days of wall clock time). Running bare metal is useful, too, but you don't have system calls there.
If you're doing fast RISC-V virtualization, like on QEMU, or doing emulation on an FPGA, you're running at >1MHz and running a "normal kernel" like Linux is totally tractable. However, it would be foolhardy to expect to jump from hardware design to immediately booting Linux.
- software emulators
- cores implemented in an FPGA
- prototype chips
Whatever it is, you only have to implement the CPU core, some RAM, and some sort of two-way communications channel -- whether a pipe, UART, USB Serial, ethernet or WIFI.
On the other end of the communications channel you run riscv-fesvr (Front End SerVeR)
pk traps systems calls, serializes the arguments, send them to fesvr. fesvr unpacks the arguments, makes the system call on the host Linux machine, serializes the results, and sends them back to the RISC-V core running in that FPGA or prototype chip or Verilator or whatever.
So your test programs get to use not only printf() but also navigate the host filesystem, open files, get the time etc etc.
A proper Linux Kernel on the system under test would require megabytes of RAM, various I/O devices etc.
pk requires only a few kb of RAM and a communications channel.