The Design of Scalar AES Instruction Set Extensions for RISC-V
eprint.iacr.org
eprint.iacr.org
Maybe I could try banging my head against the XCrypto repository this weekend.
Prime field algorithms suffer somewhat on vanilla RISC-V due to the lack of good bignum support. For example, there is no add-with-carry, nor a convenient way to get the carry-out at a low level. You have to use the set-if-less-than instruction after an addition to separately compute the carry bit.
IIUC, you're supposed to break the bignum up into (say) 56-bit chunks, and use the upper bits of the register as the carry. I'm sceptical that that works as well as it should for practical bignum applications, though. (I haven't had occasion to try it out.)
Both Chacha and Poly1305 were specifically designed for fast execution in software implementations, so IMO this approach is well-suited to those algorithms. You could easily argue that the ARMv7-M instruction "unsigned multiply accumulate accumulate long" exists solely as a bignum accelerator, for example. Perhaps fusing addition with xor is similarly worthwhile for ChaCha.
Avoiding side channels and valuing small code size can result in some really funky work arounds. For example, using a whole ECDH key exchange in a JavaCard co-processor in order to do EC point calculations. As seen in [1] slide 30.
[1]: https://www.blackhat.com/docs/us-17/thursday/us-17-Mavroudis...
I mean, x86 sets a very high bar for the sheer number of feature flags. Most ISA extensions of interest to compilers are standardized with the RISC-V Foundation †.
The mechanism is similar to other chips, and the situation so far seems to be about the same, except a couple orders of magnitude difference in the current number of flags.
† Seems it's now RISC-V International
To clarify I did not mean that this information was accessed through specialized CPU instructions directly, but that it is available through cpuinfo on Linux and such like usual.
Here's an example device tree for the HiFive Unleashed board, which describes the embedded controller as rv64imac, and the application processor cores as rv64imafdc (same as the embedded controller, but with floats and doubles): https://github.com/riscv/riscv-device-tree-doc/blob/master/e...
If your system is application-specific enough, you won't even bother with that. You'll know what extensions your SoC supports when you make your purchasing decision, and you'll know when you compile your software.
Edit: Apparently the device tree is handled by the firmware (UEFI?), so that a RISC-V OS asks the firmware to provide the device tree and I suppose if it ever gets to desktops, it'd be firmware's job to figure out what CPU is plugged into the board. The current state seems to only cater to integrated systems, though?
Re: the added question on "integrated systems": your VM hypervisor will map a device tree for VM guests as well, this is how QEMU and other hypervisors expose virtio devices to guests as well.
Sometimes it's a combo of all three.
RISC-V is a processor ISA, not a platform spec. you may be interested in the SBI: https://github.com/riscv/riscv-sbi-doc/blob/v0.2.0/riscv-sbi...
the details of how all this stuff works at a human-collaboration level has changed a bit since 1993. there's a higher level of collaboration expected. vs intel and ms bossing people around.
edit: you may also be interested in https://doc.coreboot.org/arch/riscv/index.html and https://lkml.org/lkml/2020/2/25/1209 and https://lists.gnu.org/archive/html/grub-devel/2019-02/msg000...
not exactly "desktop" yet, but hopefully you can see how that might look, given the existing linux infrastructure :)
Yes.
> After a cursory search I couldn't find anything like it.
This was the first google result for "riscv cpuid feature instruction" for me (RISC-V Instruction Set Manual v1.7):
> The mcpuid register is an XLEN-bit read-only register containing information regarding the capabilities of the CPU implementation. This register must be readable in any implementation
But IIUC it will not be accessible at user-level, so capability detection will be significantly less convenient for cross-platform library authors compared to x86.
I like the approach of hiding cpuid registers from user-space. It means, for example, the operating system can easily control what hardware features user-space will use. That's useful for debugging, for process migration, for record and replay (I maintain rr), and to "discourage" user-space from using buggy or detrimental features.
Simplifying the hardware features exposed to user-space makes evolution easier. Your OS kernel should be (needs to be) up to date, while you might have old user-space binaries that you want to keep running, so you can live with weaker compatibility guarantees for supervisor-mode-only features.
They won't work on OSes the library author didn't know about, or hasn't supplied code for, or which didn't exist when the library was created.
There are a lot of OSes.
Most likely it will be a call to libc or equivalent, and therefore be tied to libc. For generic libraries it may be the only reason for a dependency on libc.
There are even more libc variations than OSes. In Linux terms, it is almost distro-specific.
This is effectively part of the RISC-V ABI, and it means there is no OS-independent ABI.
Maybe, but right now there are only a couple OSes that run on RISC-V, where you would care to detect these things at runtime. As of now, I think FreeBSD has an AT_HWCAP, which exposes RISC-V standard extension information. (though yes, it will be some platform code, one #ifdef and a couple lines to call either getauxval or elf_aux_info).
> ...or even memcpy()
P.S. memcpy goes in libc, the same place that implements these functions, so I think probably choosing the right memcpy is not an issue.
However, I didn't see how a nonstandard extension such as this would be represented; this is concerned with representing up to 26 or so extensions, the ones with single letters assigned to them.
Same as for x86, Power, ARM, MIPS, systemz, sparc, ... ISAs: there is a flag register that the OS can query to test whether the hardware supports something.
I only know one ISA that works in a completely different way (WebAssembly), by requiring the user to attempt to perform an operation, and if that faults, then the operation is not supported. The reason WASM can do this is because it is a virtual ISA. This approach does not work well for real hardware because it makes instruction decoders larger and slower, since it prevents them from assuming that they will only be fed instructions they do know.
> This approach does not work well for real hardware because it makes instruction decoders larger and slower, since it prevents them from assuming that they will only be fed instructions they do know.
But what if you fail to do the runtime checks? And for example just feed an SSE4 binary/instructions to a processor that only supports SSE3? Then won't it segfault?
That includes trapping on an illegal instruction or something, but there is no guarantee that some CPU that was designed and built before SSE4 existed will do anything meaningful when fed with machine code that it was never intended to see.
You are basically hopping for the instruction decoder to be "really good", and not "really fast". Eg. suppose that there is only one instruction in SSE3 which has the 7th bit set. A fast instruction decoder tests that bit, and if its true, it knows exactly what instruction that is. Years later, SSE4 is designed and implemented, and it adds another instruction that also has that bit set. You are required to check whether the CPU supports SSE4, and the SSE3-only CPU will tell you that it does not. If you then go ahead and feed it a SSE4 instruction, the CPU can do really anything, including executing some other completely different instruction.
I don't see how this could be a security flaw.
AES would be non-RISCy because it won’t execute in a single clock cycle, either stopping the pipeline or necessitating complex logic.
Having said that, I don’t think RISC-V implementations are 100% RISC in that respect (is any CPU?)
Of course the line is blurry, but to me things like SIMD, floating point or AES are more like "coprocessor" extensions, it's not something that can decently be emulated using a handful of CISC opcodes. Making a quasi-religious point that CISC shouldn't support high-level instructions will doom these CPUs to irrelevance.
A CISC CPU is no less CISC if you put in on a board next to a hardware video decoder module for instance, so why should an on-die AES extension be a game changer?
The original PlayStation's CPU had a coprocessor dealing with 3D transformations (the part of the pipeline that would be handled by a vertex shader these days) so it actually supported opcodes such as "MVMVA: Multiply vector by matrix and add vector". I don't think anybody would consider that unRISC-y.