RISC-V Int. Ratifies 15 New Specs, Opening Up New RISC-V Design Possibilities
riscv.org
riscv.org
- Scalar crypto: https://github.com/riscv/riscv-crypto/releases
- Vectors: https://github.com/riscv/riscv-v-spec/releases
- Bitmanip: https://github.com/riscv/riscv-bitmanip/releases
The trouble comes when you need to share access to a memory mapped peripheral among multiple threads/processes/users etc. It can be done, but it's usually easier to manage CPU registers than peripheral devices for things like crypto operations in larger systems. Plus, you have to do access control to the peripheral (so other processes don't try and steal your key), if its all within the security boundary of a "normal" process, you get that (mostly) for free.
All of the above has caveats and exceptions, but generally (ARM, SPARC, x86, now RISC-V) take this approach.
Full GREV and my lovely GORC have also gotten lost, though the encodings for the specific REV and ORC instructions that are included are upwardly compatible with the proposed general versions.
(J Extension is about dynamic languages acceleration; stuff like code caches, and maybe providing GCs some help. I guess that's new territory so it's not as straightforward compared to say, the bitmanip extension)
The reasoning is simple. It is indeed relatively new territory. The research needs to be done, and standards should be based on solid research. This will likely take a significant amount of time.
Experimentation can be done using custom extensions, but only what's mature and proven belongs in RISC-V official extensions.
Even with more extensions ratified, that does not mean they will be available on common target hardware, thus requiring, at best, build-system gymnastics unlikely in normal software distribution. This is especially true for any not required in published profiles.
Unless you are the system architect for a massive-volume embedded application, the extensions are more like a cruel joke: "You could have had this feature if we let you, but we didn't. Peasant."
Wikipedia only lists 6 as frozen, so where did the others come from? https://en.wikipedia.org/wiki/RISC-V#Design
Updated versions of the Privileged and Unprivileged Spec PDFs will be posted to riscv.org/specifications soon.
* PMP Enhancements for memory access and execution prevention on Machine mode (Smepmp)
* RISC-V Base Cache Management Operation ISA Extensions
* RISC-V Bit-Manipulation ISA-extensions
* RISC-V Count Overflow and Mode-Based Filtering Extension
* RISC-V Cryptography Extensions Volume I: Scalar & Entropy Source Instructions
* RISC-V State Enable Extension
* RISC-V "stimecmp / vstimecmp" Extension
* RISC-V Vector Extension
* The RISC-V Instruction Set Manual Volume II: Privileged Architecture
* "Zfh" and "Zfhmin" Standard Extensions for Half-Precision Floating-Point
* "Zfinx", "Zdinx", "Zhinx", "Zhinxmin": Standard Extensions for Floating-Point in Integer RegistersThis is different from "fixed-width SIMD" which has a hard-coded vector length. To make things more challenging for the programmer/compiler, I believe most x86 SIMD versions also don't provide a "mask" register, so you're stuck with using all vector elements (AVX512 added masks).
Each has its advantages and disadvantages (esp. on the design complexity vs programmer/compiler interface complexity).
RVV also provides a mechanism to reconfigure the register file, ganging logical registers together to get longer effective vector lengths.
Who is going to write all the documentation and snippets for them? RISC-V docs seem to be mostly pdf based which isn't great.
If this is not the case then our programming model may end up being the same as it is on X86 since it's a easy subset of the functionality.
So sorry I was a bit confused because Cray/RISC-V is explicitly designed to be easy to use from high-level languages in a way that x86's SSE et al were not. So I thought maybe you had mixed up the two in your question or something. But I guess you just haven't had the pleasure of working with a Cray before!
The thing with the Intel model is that although the programming model in the abstract is probably worse (although I'm curious if it allows a wider processor), it is trivial to use conceptually if you understand roughly which instructions you want i.e. it's just a blob as far as the compiler is concerned.
The compiler I work on supports Intel SIMD, I'm not sure it could be easily made to get the most out of a vector programming model without a lot of rewrites. It could, however, basically emulate the fixed width things in terms of a vector ISA if needs be.
http://www.audentia-gestion.fr/CRAY/PDF/Cray_C_and_C___Refer...
You literally just write regular old C code doing a tight inner-loop computation, and use pragmas to tell the compiler what it needs to safely parallelize.
Of course these days you can do the same thing in any vectorizing compiler. But the point is that a modern vectorizing compiler has to do some pretty impressive transformations to generate SIMD code which looks nothing like the original, whereas the Cray code pretty much compiles to the same thing when vectorized.
In very tight situations it's common to write SIMD intrinsics directly rather than rely on the compilers ability to make the transformations itself. Intel's SIMD maybe be ugly but it is also very topologically easy to navigate, if that makes sense.
I'm going to write some arm SVE code and compare, at some point.
This is all I remember, there is probably more.
Machines with any size vector registers handle code specifying vector length of 1 (or 0!) no problem.
If you really want to make a machine with vector registers that hold only one element then that will work too, except for a handful of instructions that simply don't make sense in that case (unless you use the LMUL feature): vector permute register, slide up, slide down.
CPUs intended to run standard operating systems with shrink-wrapped software are constrained in the RVA22 profile to provide vector registers of at least 128 bits and no more than 65536 bits. But if you're doing some custom embedded custom CPU then you can make the vector registers the same size as the integer registers (32 or 64 bits). Note that if you do that, you can still usefully do vector operations on chars and shorts, and you can also set LMUL=8 to give you effectively four vector registers of 256 or 512 bits each (which might or migth not be processed serially).
Not to name drop but here's what David Patterson had to say (he's vice chair of RISC-V BoD among other things).
"One of brilliant features of RISC-v is modularity. Everyone wants an ecosystem that is adaptable but runs standard software. Defining profiles and platforms is the next thing on their slate. Binary compatibility is not the overwhelming thing in the SoC world that it was with microprocessors. Flexibility is one of the various attractive features of RISC-V."
The idea with profiles is that you create groupings of modules aimed at a specific use case.
So, yes, there needs to be some balancing of flexibility and compatibility/interoperability and there are concerns around this. (One of the processor analysts brought this up.) But people are aware and thinking about it.
ARM does something similar. They have TONS of extensions, but then group them into 8.0, 8.1, 8.2, etc then also group them with the A, R, and M designators too.
The more general bitmanip extensions contain other things useful for e.g. address arithmetic. These are somewhat orthogonal to scalar crypto.
Aside, Keller's quote is probably partly in jest. If you are in a constrained micro-controller environment something like the ZFinx extension is probably helpful beyond the "just six instructions" for code density. If you are crypto heavy, the crypto extension are going to be more helpful than "just six instructions". If your workload is parallelisable and regular, vectorisation helps you more than "just six instructions" and so on.
One size doesn't fit all.
All of the discussed extensions helps specific workloads, but unless your workload is, say, 100% encryption all the time, then the crypto extension will only provide a trivial improvement on the _overall_ performance.
Vector is a little bit different, but it (like AVX2/512) comes at a very significant cost and you better have software that can take advantage of it.
If there is really a significant win for a certain type of server workloads, that community will make its own profile and hopefully be able to get chips that utilize that.
The problem is that there are also many mixed workloads and having lots of general compute can work pretty well if you want to run a broad set of extinctions.
RISC-V is sort of a fluid spectrum from highly specialized to highly general depending on the use case.
True, but a standard that is too malleable isn't really a standard at all.
RVA22[0] is the first such profile, and among other important things which go a long way to ease cross-vendor software compatibility, it does require RVA22U and RVA22S, which in turn require a set of extensions.
[0]: https://github.com/riscv/riscv-platform-specs/blob/main/risc...
The idea is that what is dominates is software. If you add your own extensions, literally all software in the world wont support it. You will need to provide a huge amount of stuff to fully take advantage of that.
The availability of software both open and commercial on top of standardized profiles targets should be what manufacturers target.
Early on of course, manufactures have provided things that are not standard yet. However over time, does it really make sense to supply your own bit manipulation extension? As the standard grows the waste majority of application should not require or be really improved by proprietary extensions.
Of course if somebody comes along and makes a chip that is just vastly better then what anybody else has with some extensions. That could break that paradigm and people might embrace it.
You just choose whether to use that version of the library or another one that uses normal instructions.
It's no exaggeration to say that many of those extension instructions might exist in only one function in one library on your entire Linux (or Android, FreeBSD, whatever) system.
To some extent the Vector extension can be like that. For most programs they'll just pick up vectorised versions of memcpy, strlen and so forth. In other programs (generally ones you compile yourself) you might want to use the vector extension directly -- maybe with auto-vectorisation in time. LLVM can do a bit of that already.
Only a few of the extensions have instructions that can profitably weave their way into every part of your code. The Bitmanip extension is like that. You really want to know whether your target processor has B or not.
RISCV offers lots of official extensions to choose from, such as M, A, F, D, P, V, .... In addition you have the 32 vs 64 bit data width parameter. Any specific ISA will have to instantiate those parameters, like e.g. so: RISCV32MFP or RISCV64MAF. Any implementation of e.g. RISCV64MAF will have to implement in silicon exactly those assembly command (and supporting features) that the M, A and F extension demand, with 64 bit register width.
Like in OO-programming the class constructors take arguments that parameterise the created object.
------
Regarding an implementation, given that RISCV is an ISA, not an ISA implementation, you need to provide a functional model. The official standard is [1] but it's a bit behind the ratified extensions. For example [2] defines the (ISA-visible)
registers, while [3] gives you the instruction decoding and execution clause for the most base instruction set. [4] describes part of one of the available address translation modes (for the 32 bit variant of the ISA). Note: in modern processors page-table walks are hardware accelerated, so OS and processor need to use the same format here, which is why this is part of the ISA.[1] https://github.com/riscv/sail-riscv/tree/master/model
[2] https://github.com/riscv/sail-riscv/blob/master/model/riscv_...
[3] https://github.com/riscv/sail-riscv/blob/master/model/riscv_...
[4] https://github.com/riscv/sail-riscv/blob/master/model/riscv_...
So the major Linux distros agree on a set of instructions and that's called a profile. Same for embedded and others eventually.
You can add your own extensions for yourself if you want. You can also make extentions and try to make it a sudo standard. Or you can attempt to make it into a standard extention.
To be a standard extension it has to go threw a long process and it will likely be tapped out multiple times before it is ever ratified. Once its ratified it will find its way into profiles.
So for example standard Linux distros now use RV64GC, likely the next version of the Linux profile will include more of the new instructions.
But yes, the goal is not to create a 'universal binary'. But a reasonable compromise between reuse and specialization.
Yes. Just like you cannot run Pentium code on a 386 because they added new extensions. Or how Scheme isn't really a programming language but more like a _family_ of very nearly compatible languagues. RISCV has multiple targets and so so they have very different needs from embedded automotive to desktop. But with a common core is easier to develop and share tooling.
On the programming side, you can detect at runtime feature support and use specific code path accordingly, or decide at compile time that you require a specific CPU feature and then your binary will just not work on CPUs without the feature.
* RISC-V Vector instructions seem like a huge win for all forms of HPC. x86 is getting vector instructions & the wins have been immense. Rather than a wide range of specific SIMD instructions, vector instructions seem like a far more general & easier to scale up & down implementation strategy. Not everyone has to implement!
* RISC-V Hypervisor specifications seem required for modern computing, where VM's are commonplace. Have to have this specification. Not everyone has to implement!
* RISC-V Scalar Cryptography specifications providing accelorated cryptography seems like another have to have modern in data-centers.
Worth re-iterating what's been said already: extensions are just that: extensions. They're not required. I'm not sure what the current state is, of code detecting & use the accelerated implementation when available, using soft-fallbacks otherwise. For things like cryptography, usually it's a library, openssl or someone, where the library is the reference implementation, with special paths written in for using harware where available.
All three of these are complex enough to definitively increase the design/verification time for any core that implements them, though, that's for sure. (A net effect of this is that while there are tons of simple in-order cores, actual "production" RISC-V cores with features like this will remain rare...)