BTW I've been writing a document on the topic of how Linux distros will support hardware with and without compressed: https://github.com/rwmjones/rhrq-riscv-extensions/blob/main/...
Firstly, as you probably know the Clang/LLVM model is quite different to gcc/binutils in terms of build-time configuration. Any clang can target any architecture (unless it was disabled at build time) using --target=$TRIPLE (see <https://clang.llvm.org/docs/CrossCompilation.html#target-tri...>). In practice it's not as useful as it sounds of course because you need appropriate sysroots. You can control the default with LLVM_DEFAULT_TARGET_TRIPLE, but without patching clang it's going to enabled compressed for the riscv64-unknown-linux-gnu triple. I think the standard way of controlling other default flags would be to deploy a configuration file <https://clang.llvm.org/docs/UsersManual.html#configuration-f...> - though I'd need to check which logic took precedent if you explicitly passed --target=riscv64-unknown-linux-gnu. Ultimately we'd need to change the meaning of the current linux triples, or decide upon new ones I think.
Thanks for sharing that doc - really helpful. I don't know if it's different with GNU as, but not that if a file contains `.option rvc` that _isn't_ sufficient to set the EF_RISCV_RVC ELF flag with LLVM. Additionally, I'd imagine in general that people using .option in inline asm may see different behaviour for Clang vs GCC. Clang doesn't generate assembly output that then gets passed to the assembly - ELF emission happens directly (of course, with the assembler being invoked as needed for inline asm blocks), and so it's rather difficult to replicate the effect you'd see with GCC generating a .s file that's then assembled.
On the flip side, Minimax is a RISC-V design that only implements compressed instructions: https://news.ycombinator.com/item?id=33422717
EDIT: Looks like I'm not alone: https://news.ycombinator.com/item?id=29667135
He is completely right and I have had those exact problems: Not being able to figure out where you are in the instruction stream annoys me when it comes to RVC. I actually dislike it so much that I intentionally disregard it in my RISC-V emulator and focus on other things, eg. RV64GVB etc. I support C-extension, but it's hard to make it run fast. I wish they solved compression another way. We have very fast compressors/decompressors that can use dedicated dictionaries that could have been used in ELF loading instead.
I managed to find this PDF that does admit that it has a modest 2-3% benefit whenever the code wouldn't normally fit in icache: https://forums.macrumors.com/attachments/a-case-to-remove-th...
The real issues are two-fold: Compressed instructions consume 75% of the instruction encoding space. The thought is that without compressed instructions occupying that space, more 32-bit extensions can be added before we need to use longer instructions.
Secondly verifying designs which use unaligned instructions is difficult because there are many different cases to consider (especially instructions crossing cache lines and pages). The verification thing comes from a particular vendor who already design very high end chips - you are likely to have several in your possession - and I have to assume they know what they're talking about.
Is this a practical concern? You can trap the case where an insn is potentially spanning two pages and handle it in software. As for cache lines, these are small enough that the issue could come up regardless, e.g. from a vector load insn, which is something you do want to support.
I don't see how this applies to RISC-V compressed instructions? As far as I understand they're always 16-bit and always come in pairs such that 32-bit instructions are still perfectly aligned.. no?
I'm not entirely convinced by the "huge i-cache" argument either. Fitting more instructions into the instruction cache is always a benefit. Moving less data around is always a benefit. And the RISC-V compressed instruction scheme is so simple it doesn't really cost much. Put another way, if they supported compressed instructions maybe they wouldn't need such big I-caches and could fit another core or accelerators instead?
> Compressed instructions consume 75% of the instruction encoding space. The thought is that without compressed instructions occupying that space, more 32-bit extensions can be added before we need to use longer instructions.
This argument I can buy though. I could see the point that there are more complex instructions that they want to use the encoding space for in a high end server architecture. But can they? Is that supported by the standard?
Nope, 32-bit insns are naturally aligned if you don't use C but only 16-bit aligned if the C extension is present. So a single insn can indeed cross a cacheline or page boundary. This would be fixed if C insns "always came in pairs" (at the expense of some extra C.NOP's) but they don't. (Though this could be made an optional extension of its own.)
> But can they? Is that supported by the standard?
C has always been an optional extension, and when not using it you free up that part of the encoding space. Additionally, the 3 blocks of encoding space that make up C are largely self-contained so the extension itself could be split up, which would allow for freeing plenty of space while still keeping most of the benefit of compression.
The ISA manual hints at this: with regular instructions allowed to be 16b aligned, regular & C instructions can be mixed freely. This also helps improve the code density.
Requiring C instructions to come in pairs, reduces that advantage.
So it's a kind of all-or-nothing affair.
Probably RISC-V designers ran a bunch of code simulations, and concluded that "mix regular & C instructions freely" weighed heavier than "simplify RISC-V cpu designs that support C subset".
Note that cpu design effort is a 1-time cost (when designing new cpu). Whereas adding spurious NOPs would be a recurring cost for all software using the C subset.
How much space is spare after the rest of the (mainstream) extensions are included?
Yes, compressed instructions rely on variable size opcodes.
No, it isn't like x86 where it's really complex for the decoder to know where instructions start and end. It has been designed to avoid that.
As for actual implementations, every server chip that's been announced so far supports them.
This includes Tenstorrent Ascalon, the one with the 8-wide decoder.
Higher code density means more code fits in cache, or less cache is needed (and can thus be clocked higher). It helps very high performance implementations.