Some Criticisms of RISC-V
gist.github.com
gist.github.com
I'd also emphasise the lack of indexed load/store instructions again. It is a disaster for any sort of array indexing where you're addressing >1 array using the same index. I found this in the context of multi-precision arithmetic for crypto, but examples come from all over.
On the author's point about multiply and divide in the same extension: crypto is another good example. Lots of crypto really benefits from multiply, but doesn't need divide.
The most important thing about RISC-V is the idea behind it's openness as a standard and the ecosystem around that standard. The engineering of the ISA itself is not what makes it remarkable, and actually leaves a lot to be desired.
Most software doesn't use rotate operations. When you do need one, it takes three instructions to synthesize it if the shift count is a constant, or four otherwise. Unless you're doing nothing but rotates you're not going to notice it. If you're doing any memory loads that miss in the L1 cache to get the data you're rotating then you're also not going to notice it.
The same goes for the indexed loads and stores. Measure it, don't just go on your feels.
The main problem with the the original post is that the author criticizes things based on his sense of aesthetics, and what he's used to, not on any kind of actual scientific experimentation and measurement of different options.
The whole premise of H&P's "Computer Architecture: A Quantitative Approach" is that you should base decisions about what to put into your hardware and what to leave out based on actual data, not on what seems prettier or more orthogonal to you.
https://www.theregister.co.uk/2019/07/27/alibaba_risc_v_chip...
"The aforementioned RV64GCV designation means the Xuantie 910 implements the base 64-bit RISC-V ISA (RV64G), supports compact 16-bit-wide instructions (C) as well as the usual 32-bit-wide instructions, and supports still-in-development vector math operations (V). Interestingly enough, though, it is also said to include 50 non-official instructions for assisting and accelerating various tasks, from memory management and CPU core wrangling to storage access."
I think this is elegant, at least superficially. However this ambiguity makes tracking the callstack harder, which RISC-V solves by introducing the convention that the x1 register is used for the return address. Which means that the normally irrelevant choice of register suddenly has performance implications, which is not so elegant.
It takes three arguments: rd, the destination register to put the return address in (it should always be the ABI-defined link register, e.g. t1), rs1, the source register that jump should be relative to (it should always be the link register for a return, and 0 for a call/branch). To a software guy, there's elegance in unifying the three cases, but there's really no such principle in hardware. It just makes the implementation more complex, since branch/call/return all greatly differ for prediction, prefetching, etc.
I was astonished that they didn't consider bitwise anything -- including rotate and popcount -- worth bothering about. It is sad they settled on 1, not ~0 for their bool true value, in the ABI. Do they even have short load-immediate -1, 0, and 1 instructions?
Micro-op fusion is great, but only if you can keep your instruction decoder sated.
I also remember a friend that worked at a failed processor company in the 1990's. Said the real reason the company went under was the academic designers thought programs 'do integer math' when modern programs 'do string manipulation'.
What do you believe is the reason that ARM does no market this acronym anymore? ;-)
(Obviously not saying OP is an uninformed pundit — topic just sparked a funny memory)