RISC-V Announces Ratification of the RVA23 Profile
riscv.org
riscv.org
RISC-V has taken an approach where they have a relatively small core instruction set[1], and then a relatively large set of instruction set extensions.
Think of it like SSE and AVX on the x86 platform, turned up to 11.
This makes it more attractive to chip designers, as they can pick and chose what to implement, making fairly specialized chips at potentially reduced cost.
However this also makes it more difficult to target for programmers, as technically a RISC-V chip can have any combination of these extensions.
To make things a bit more predictable they've come up with these profiles, which is just the core plus a defined set of extensions.
So if you find a chip that says it follows the announced RVA23 profile, you know for example the vector instructions must be available.
Is there any way to distribute a binary that uses "core" and then have users recompile the binary to use their local extensions?
That's completely fine and standard. It's already what people do with x86 too, if you want to ship say AVX-512 code you have to deal with the fact that the latest Intel client CPUs don't yet have support (product segmentation issues), so at runtime you have to check if the extension is supported.
The situation is not different with RISC-V. You compile your binary for the standard profile, and optionally if you want to use a specialized extension to optimize some hot loop, you check at runtime if it's available.
Shirley you're joking? Did you miss APX, coming soon in Panther Cove?
- expands x86_64 from 16 to 32 GPRs
- adds 3-address instructions
- updating flags is optional
- push/pop 2 registers with one instruction
- more conditional execution instructions, including loads & stores
GCC 14 already supports 32 GPRs, 3-address, push2/pop2. GCC 15 (head) supports NF (No Flags).
Yes[1].
> Do you need annual recompilations of software..?
The profiles is designed to provide guarantees of what's available. Just because RVA23 was released now, does not mean the RVA22 CPUs and the software that runs on those CPUs stop working.
Also, at least for the foreseeable future, the profiles will be extensions of the previous ones[2]. Thus RVA22 should be a subset of RVA23, and any software targetting RVA22 should run fine on a CPU supporting RVA23.
[1]: https://github.com/riscv/riscv-profiles/blob/main/src/profil...
[2]: https://github.com/riscv/riscv-profiles/blob/main/src/rva-pr...
I feel at some point you're either going to have penalties from the all the platform check and jumps or you're going to only use these extension in critical hotpaths and they get underutilized (what you see with SSE in x64 and the whole mess on ARM)
I'll agree it sounds like it could get messy. That said, I'm sort of imagining the core set of extensions to not grow significantly larger. Perhaps they'll add some granularity to the various profile levels, like embedded application CPUs like in set-top boxes might not need all the features of a full-blow desktop or server CPU.
Right now there is RVA20 (almost all shipping hardware), RVA22 (only Canaan K230 and SpacemiT K1/M1 shipping now, more coming next year e.g. SG2380), and RVA23 just ratified, hardware in probably 2-3 years.
Worrying about multiplication of RISC-V profiles is ignoring the existing reality of every other platform.
There are currently 17 different releases of ARMv8-A and ARMv9-A.
Have you seen the extensions list on any recent x86? My laptop reports:
fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf tsc_known_freq pni pclmulqdq dtes64 monitor ds_cpl vmx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb ssbd ibrs ibpb stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid rdseed adx smap clflushopt clwb intel_pt sha_ni xsaveopt xsavec xgetbv1 xsaves split_lock_detect user_shstk avx_vnni dtherm ida arat pln pts hwp hwp_notify hwp_act_window hwp_epp hwp_pkg_req hfi vnmi umip pku ospke waitpkg gfni vaes vpclmulqdq rdpid movdiri movdir64b fsrm md_clear serialize arch_lbr ibt flush_l1d arch_capabilities vnmi preemption_timer posted_intr invvpid ept_x_only ept_ad ept_1gb flexpriority apicv tsc_offset vtpr mtf vapic ept vpid unrestricted_guest vapic_reg vid ple shadow_vmcs ept_mode_based_exec tsc_scaling usr_wait_pause
So the world is stuck with x86-64v3, more or less haswell/zen1 level.
The only things that have mix-and-match extensions are embedded devices where you choose exactly what chip you buy (or build), you know what extensions it has, and you compile for that set of extensions. There is no confusion because everything is under the control of one company.
Machines intended to run software distributed in binary form follow the RVA specs which are few in number, only added to every few years, and each one includes everything in all the previous ones.
There is again no confusion -- or at least no more than in any other ecosystem that is not permanently frozen.
RVA23 is pretty much equivalent to ARMv9 or the latest x86_64 spec with a variable length version of AVX-512.
However, that wouldn't mean every core must implement every possible instruction - instead, unimplemented features would be emulated. The standard would require that hardware trap everything unimplemented, and that software always contain a fallback implementation.
Then, the fragmentation problem turns from a "it won't work" problem into a "it might not be quite as fast" problem - and for general users, that effectively means everything is 100% compatible.
RISC-V chose a different approach, because this one, quite frankly, would not have worked.
For (fixed-length-)vector extensions in particular, I’m not sure that would work well. For nontrivially vectorizable things (i.e. you’re not just doing a bunch of math in parallel), the usual tradeoff is that you use a couple times more compute, but because the vector unit can provide a couple dozen or more times the compute per clock the result is faster despite the apparent waste. Emulating such code, on the other hand, will likely be miserable.
Note that both LLVM and GCC will try to auto-vectorize code and emit vector instructions.
It's much better to compile multiple versions of the code or multiple binaries.
Also, any OS kernel could do transparent emulation like that with no need for CPU assistance (beyond trapping on unsupported instructions, which all modern ISAs of course already do), so it's more of a Linux/Windows ABI issue. You can also write an LD_PRELOAD library that does the emulation.
This is no different from how x86 works today. Both Windows and Linux OSes and their binaries assume some minimum level of x86-64 support (eg. x86_64-v3 will be the base "profile" for RHEL 10 [1]). For other things like AVX-512, portable binaries test for the "extension" before using.
[1] https://developers.redhat.com/articles/2024/01/02/exploring-...
- https://riscv.org/about/faq/ j - https://riscv.org/membership/ (I thought this was going to be expensive, but 2k USD/yr -> 5kUSD/yr after 2yr can keep a small company going)
Welcome to non-copyleft hardware designs.
It's how things work anyway, things that are on the edge of what can be done are good for running business and maybe not finalized enough to become standards.
Secondly, extensions cover a lot of things. They can do much more than just introduce instructions, like introduce new hardware registers. The RVA23 profile that's under discussion here for example includes a hypervisor extension[1]. I don't see how you can emulate that in software after the fact in a meaningful way.
Third, RISC-V was designed to be used all the way from microcontrollers to application processors. Especially for the embedded use-cases, you're targeting a well-defined collection of processors, so which extensions you can use is well-known and not subject to changes willy-nilly.
The RVA23 profile that's just been ratified is aimed at more general processors, which will run more varied software. There it makes sense to agree on a well-defined subset of extensions, which is why they do exactly that through these profiles.
In fact the hypervisor extension was explicitly designed to be relatively efficient to implement in software emulation.
It even says so at the URL you failed to include:
"The hypervisor extension has been designed to be efficiently emulable on platforms that do not implement the extension, by running the hypervisor in S-mode and trapping into M-mode for hypervisor CSR accesses and to maintain shadow page tables. The majority of CSR accesses for type-2 hypervisors are valid S-mode accesses so need not be trapped. Hypervisors can support nested virtualization analogously."
https://five-embeddev.com/riscv-priv-isa-manual/Priv-v1.12/h...
Fair point, I stand corrected on that one. Should have checked more thoroughly before posting.
> It even says so at the URL you failed to include
That site doesn't render right on my mobile, which I was on at the time, so couldn't read the contents.
> "[...] by running the hypervisor in S-mode [...]"
That doesn't sound very efficient if all you got is an M-mode-only chip though.
In any event, you also have extensions like Ztso[2], which even the spec says to just let the binary crash in event of non-presence, or Zam[3] which you technically could emulate but surely would be dog slow.
[1]: https://five-embeddev.com/riscv-priv-isa-manual/Priv-v1.12/p...
[2]: https://five-embeddev.com/riscv-user-isa-manual/Priv-v1.12/z...
[3]: https://five-embeddev.com/riscv-user-isa-manual/Priv-v1.12/z...
[4}: https://www.riscfive.com/2022/12/07/sifive-intelligence-fami...
There is approximately a light-year of space between an M-mode-only chip and anything where you'd consider running a hypervisor.
Suggesting that you want a hypervisor is implicitly saying that you already have S and U modes, virtual memory, page tables, MMU etc.
Otherwise you're at much the same point as those news stories about how someone is running RISC-V Linux on an AVR or Pi Pico (original) or vim macros, by writing a RISC-V emulator on one of those.
RISCV is actually better than most in this. First of all there is a core ISA that will run on everything and the extensions can be setup to be ignored when running on a system without them, defaulting to some core-ISA alternative.
The only alternative is to get the entire ISA spec done before letting anyone make anything... and that's just not happening.
The hardware would provide some assistance for said fallback mechanism - for example in the form of a pointer to a "fallback vector". If you execute unknown instruction 0x80 then code at the fallback pointer+0x80 will be executed instead.
The binary doesn't have to use this fallback mechanism - for performance critical binaries which want to run fast on different hardware implementations, the compiler could use feature detection to make multiple implementations of some functions, the way many SSE/AVX extensions are dealt with in the x86 world.
Binaries that aren't intended to run on bare metal, but instead run on some OS (eg. windows/linux) could instead specify as part of their ABI which instructions are part of the core, and the OS will ensure that anything unsupported by the hardware but required by the binary will be provided by the OS.
Your comment about the compiler using feature suggestions has got nothing to do with RISCV, it's up to the compiler developers how to do that and if they want to. This is how it works right now. The compiler inspects the hardwarew and works accordingly.
RISCV dont get to decide how the OS works. It's up to the OS developers on how a binary is handled. They can't stop me creating an OS that doesn't filter unsupported instructions and therefore the guarantee cannot be kept.
You could have an interrupt style thing for unsupported instructions with suggestions on how software/the OS handle them, but whether or not they choose to handle them is not RISC-Vs problem.
That is exactly what RVA23 does. It specifies what must work. It doesn't specify the performance of any instruction -- that is between a hardware vendor and their customers as to which ones are fast.
> If you execute unknown instruction 0x80 then code at the fallback pointer+0x80 will be executed instead.
RISC-V instructions are 4 bytes long (i.e. 2^32 of them), not 1 byte. Using that simple technique, with just a pointer to the actual handler included in the table, the table will be 32 GB in size.
Obviously impossible, given that commercially-available RISC-V chips start from 2 KB RAM, 16 registers, and 48 MHz.
> wouldn't mean every core must implement every possible instruction - instead, unimplemented features would be emulated
And that is exactly what RVA23 is.
Every instruction must work. Nothing requires them to be fast. Missing instructions or capabilities (e.g. misaligned load/store) are handled in M-mode software, transparent to both the User program and the operating system.
> Then, the fragmentation problem turns from a "it won't work" problem into a "it might not be quite as fast" problem - and for general users, that effectively means everything is 100% compatible.
That is exactly the situation with RVA23.
Secondly, that's like asking cpu to do the job of a compiler. The compiler has one advantage over a cpu: it runs in a completely situation where memory and time do not matter. The cpu doesn't have the luxury of trading off memory or time.
[1] I wrote a quite detailed paper about RISC-V extensions last year: https://research.redhat.com/blog/article/risc-v-extensions-w...
I believe I was reworking the instruction decoder on my own core and the specification sheet left some open questions
I would say:
A) Numbers in instruction set extensions tend to indicate the bit width, thus my confusion.
B) If you do later want to introduce wider instructions, its going to be confusing.
C) Software moved away from whole integer numbering for a reason. is RVA25 a completely new instruction set, a bug fix, or a superset of RVA23? RVA 1.1 or RVA 2.0 gives you more of a clue as to what youre dealing with.
D) Come 2100 and RVA00 we are going to have numerous issues with software checking that 'RVA >= 23'. I would like this one to be humourous, but unfortunately, experience shows it probably wont be.
C,D) The current naming scheme is RVAxx where xx is increased for every major profile update, one that adds new mandatory extensions. Minor releases are RVAxx.y (iirc) which don't change mandated features but may allow more optional extension. The profiles are supposed to be backwards compatible, and will have a slow release cadence. The increment isn't fixed, but it's still unlikely we'll run out of two digit names. Regardless, if we ever are at RVA70, it's trivial to preserve ordering by going to something like RVA710 next.
I wish it to be successful. Worst case scenario, RISC-V is a nice byte code, much less toxic than any higher level computer language out there.
They have to be very careful about what they put in RVxx profiles: advanced application developers know already they will have to query the hardware in order to install proper (this must not be part of any file format to avoid toxic complexity: it must stay in full control of the application) machine code.
Currently I do personally code my "core" RISC-V own little applications, which I do interpret on x86_64/linux and using another executable file format than elf (of my own) but with transparent binary compatibility (aka no need to patch the kernel).
And guess what, to be not dependent on those horrible compilers like gcc/clang or ultra complex syntax languages (all full of planned obsolesence even on the medium run), or those completely PI locked hardware ISAs... is just a breeze of fresh air.
Ofc, this is a compromise as it is impossible to run a reasonable (interpretation may vary a LOT here) linux desktop without some of the worst software or dependency out there... but this is moving forwards, and it feels good.
RISC-V is not byte code though, it was not designed for efficient execution by a software interpreter.
And by what metric would it be less toxic than any higher level computer language out there?
> Bytecode (also called portable code or p-code) is a form of instruction set designed for efficient execution by a software interpreter.
It's not bytecode.
It is if a byte is 16 bits.
> it was not designed for efficient execution by a software interpreter.
It's not bad at all for that. On my i9-13900HX laptop, if I take http://hoult.org/primes.txt and compile it for x86 and for RISC-V:
- 5.6 seconds for riscv64-linux-gnu-gcc -O primes.c -o prime run in QEMU [1]
- 3.8 seconds for gcc primes.c -o primes run natively
- 2.0 seconds for gcc -O primes.c -o primes run natively
Emulating RISC-V is worse than forgetting to compile with `-O`, but not much. It's far faster than Python or Ruby and comparable to Java, C#, JavaScript, or WebASM.
Also, QEMU is not the fastest emulator around, just the most flexible and complete. The experimental RV8 and the nearly-ready-for-prime-time RVVM [2] are much closer to native speeds.
Note that if I do exactly the same thing for arm64 (i.e. docker/QEMU) the primes program takes 14.4 seconds and the whole Ubuntu in docker experience just feels a lot more laggy. Running the x86_864 binary in qemu-x86_64 instead of natively takes 10.5 seconds.
RISC-V is a lot easier to emulate quickly than Aarch64 or x86.
[1] or if you have Docker Desktop (on Mac, Windows, or Linux) then you can just work like a native:
bruce@i9:~$ docker run --platform linux/riscv64 -it riscv64/ubuntu
Unable to find image 'riscv64/ubuntu:latest' locally
latest: Pulling from riscv64/ubuntu
53300d777b1a: Pull complete
Digest: sha256:6a392b2c64f4e0021bfcff62e58975ddce0f1eccc5a72708b114aeb50999ff22
Status: Downloaded newer image for riscv64/ubuntu:latest
root@5f2edc942403:/# apt update
:
root@5f2edc942403:/# apt install wget gcc
:
root@5f2edc942403:/# wget -q http://hoult.org/primes.txt
root@5f2edc942403:/# mv primes.txt primes.c
root@5f2edc942403:/# gcc -O primes.c -o primes
root@5f2edc942403:/# ./primes
Starting run
3713160 primes found in 5609 ms
216 bytes of code in countPrimes()
root@5f2edc942403:/#
[2] https://github.com/LekKit/RVVM> Bytecode (also called portable code or p-code) is a form of instruction set designed for efficient execution by a software interpreter.
It's not Bytecode.
I would define bytecode as:
- Closer to machine operations than something like an AST, which is where I think the "efficient execution" part of the wikipedia definition came from
- Not directly executable on any physical CPU, or at least not typically. There are some CPUs that can execute Java/Python bytecode (sort of) but they're in the minority.
While not strictly required, bytecode tends to have opcodes meant to be used with a virtual machine, such as dealing with objects instead of linear memory.
This is all to say, the proper terminology for RISC-V is an ISA, but the distinction is more about how it's being used than how it's designed.