VRoom A high end RISC-V implementation
moonbaseotago.github.io
moonbaseotago.github.io
An example of his crazy coding chops, he was frustrated by the lack of verilog licenses at the place he worked back in the early 90s. His solution was to whip up a compliant verilog simulator, then wrote a screen saver that would pick up verification tasks from a pending queue. They had many macs around the office that were powered 24/7, and they could chew through a lot of work during the 16 hours a day when nobody was sitting in front of them. When someone sat down at their computer in the morning or came back from lunch, the screen saver would just abandon the simulation job it was running and that job would go back to the queue of work waiting to be completed.
I was designing Mac graphics accelerators I'd built it on a some similar infrastructure I'd built to capture trace from people's machines to try and figure out where QuickDraw was really spending it's time - we ended up with a minimilistic graphics accelerator that beat the pants off of everyone else
So in the end I open sourced the compiler ('vcomp') but it didn't take off
I noticed that modern CPUs are optimized for legacy monolith OS kernels like Linux or Windows. But having a large, multimegabyte kernel is a bad idea from a security standpoint. A single mistake or intentional error in some rarely used component (like a temperature sensor driver) can get attacker full access to the system. Again, an error in any part of the monolith kernel can cause system failure. And Linux kernel doesn't even use static analysis to find bugs! It is obvious that using microkernels could solve many of the issues above.
But microkernels tend to have poor performance. One of the reasons for this could be high context switch latency. CPUs with high context switch latency are only good for legacy OSes and not ready for better future kernels. Therefore, either we will find a way to make context switches fast or we will have to stay with large, insecure kernels full of vulnerabilities.
So I was thinking what could be done here. For example, one thing that could be improved is to get rid of address space switch. It causes flushes of various caches and it hurts performance. Instead, we could always use the single mapping from virtual to physical addresses, but allocate each process different virtual address range. To implement this, we could add two registers, which would hold minumum and maximum accessible virtual addresses. It should be easy to check the address against them to prevent speculative out of bounds memory accesses.
By the way, 32-bit x86 architecture had segments, that could be used to divide single address space between processes.
Another thing that can take time is saving/restoring registers on context switch. One way to solve the problem could be to use multiple banks (say, 64 banks) of registers that can be quickly switched, another way would be to zero out registers on return from kernel and let processes save them if they need it.
Or am I wrong somewhere and fast context switches cannot be implemented this way?
VRoom! largely has physically tagged caches so they don't need to be flushed, the BTC is virtually tagged, but split into kernel and user caches, you need to flush the user one on on a context switch (or both on a VM switch) - also the trace cache (L0 icache) will also be virtually tagged. VRoom! also doesn't do speculative accesses past the TLBs.
Honestly saving and restoring kernel context is small compared to the time spent in the kernel (and I've spent much of the past year looking at how this works in depth).
Practically you have to design stuff to an architecture (like RISCV) so that one can leverage off of the work of others (compilers, libraries, kernels) adding specialised stuff that would (in this case) get in to a critical timing path is something that one has to consider very carefully - b ut that's a lot of what RISCV is about - you can go and knock up that chip yourself on an FPGA and start trialing it on your microkernel
ASID = Address Space Identifier. It's a tag that uniquely identifies each processes' entries in the TLB. This ensures that your TLB lookups can be limited to the valid entries for the process, so you don't need to flush the TLB on context switch.
One thing I've done in VRoom! which is an extension on to the RISCV spec is that if we have an N hart SMP CPU (for example a 2 cpu SMT system) we use log2(N) bits of the ASID to select which hart/cpu a TLB entry belongs to - from a programmer's point of view the ASID just looks smaller.
However there's a VRoom! specific config bit (by default off) that you can set if you know that the ASIDs you are going to use for all your CPU's effectively see the same address space - if you set that bit then the per-cpu portion of the ASID tags (in the TLB) become available (ie to the programmer the ASID looks bigger) - it's a great hack because it doesn't get into any critical paths anywhere
I think a few other ARM customers were intrigued by the security possibilities, but the vast majority were more like “what is this bizarre thing, I just want to run Unix”, so the feature disappeared eventually.
Here’s some ARM documentation if you want to pull this thread: https://developer.arm.com/documentation/dui0056'/latest/'cac...
PalmOS was another one that worked similarly. https://www.fuw.edu.pl/~michalj/palmos/Memory.html
* Single 64-bit address space. Caches use virtual addresses.
* Because of that, the TLB is moved after the last level cache, so it's not on the critical path.
* There's instead a PLB (protection lookaside buffer), which can be searched in parallel with cache lookup. (Technically, there's three: two instruction PLBs and one data PLB.)
I'd also like to see further work related to this: https://core.ac.uk/reader/161119546
There's been a lot of recapitulation and growth in the language space recently as well, showing up in languages like Zig and Rust, paving the way for better utilization across heterogeneous and many core architectures. I feel like Rust's memory semantics don't hurt the mill either, and may help a lot.
The various variants of L4 have pretty good context-switch latency even on traditional CPUs, and seL4 in particular is formally proven correct on a few platforms. Spectre+Meltdown mitigation was painful for them, but they're still pretty good.
Lots of microcontrollers have no MMUs but do have MPUs to keep a user task from cabbaging the memory of the kernel or other tasks. Not sure if any of them use the PDP-11-style base+offset segment scheme you're describing to define the memory regions.
Protected-memory multitasking on a multicore system doesn't need to involve context switches, especially with per-core memory.
Even on Linux, context switches are cheap when your memory map is small. httpdito normally has five pages mapped and takes about 100 microseconds (on a 2.8GHz amd64 laptop) to fork, serve a request, and exit. I think I've measured context switches a lot faster than that between two existing processes.
Multiple register banks for context switching go back to the CDC 6600's peripheral processor (FEP) or maybe the TX-0 on which Sutherland wrote SKETCHPAD; it has a lot of advantages beyond potentially cheaper IPC. Register bank switching for interrupt handling was one of the major features the Z80 had over the 8080 (you cn think of the interrupt handler as being the kernel). The Tera MTA in the 01990s was at least widely talked about if not widely imitated. Switching register sets is how "SMT" works and also sort of how GPUs work. And today Padauk's "FPPA" microcontrollers (starting around 12 cents IIRC) use register bank switching to get much lower I/O latency than competing microcontrollers that must take an interrupt and halt background processing until I/O is complete.
Another alternative approach to memory protection is to do it in software, like Java, Oberon, and Smalltalk do, and Liedtke's EUMEL did; then an IPC can be just an ordinary function call. Side-channel leaks like Spectre seem harder to plug in that scenario. GC may make fault isolation difficult in such an environment, particularly with regard to performance bugs that make real-time tasks miss deadlines, and possibly Rust-style memory ownership could help there.
As I understand, in microkernel OSes most system calls are simply IPCs - for example, network card driver passes incoming packet to the firewall. So there is almost no kernel work except for context switch. That's why it has to be as fast as possible and resemble a normal function call, maybe even without invoking the kernel at all. Maybe something like Intel's call gate, but fast.
> they aren't compatible with anything that calls fork().
I wouldn't miss it; for example, Windows works fine without it.
I think this means we won't be seeing 'call gate' equivalents that perform close to subroutine calls on high end systems any time soon if at all
And specialized versions of this principle predate computers: a walkie-talkie has the privilege to listen to sounds in its environment, a privilege it only exercises when its talk button is pressed and which it does not delegate to other walkie-talkies, and the communication latency between two such walkie-talkies may be tens of nanoseconds, though audio communication doesn't really benefit from such short latencies. The latency across a SATA link is subnanosecond, which is useful, and neither end trusts the other.
In this case your data is likely traversing the memory hierarchy far enough so that the message data gets shared (more likely the sending data goes into the sending CPU's data cache and the receiving one will use the cache coherency protocol to pull it from there) - that's likely to take of the order of a pipe flush to happen.
You could also have bespoke pipe-like hardware - that's going to be a fixed resource that will require management/flow control/etc if it's going to be a general facility
Also note that stealing a cache line can be very expensive, if the CPUs are both SMT with each other it's in the same L1, almost 0 cost, if they are on the same die it will be a few (4-5?) clocks across the L2/cache coherency fabric but if they are on separate chiplets connected via a memory controller with L3/L4 in it then it's 4 chip boundary crossings - an order or 2 in magnitude in cost
Multithreading within a security boundary is one way to "synchronously wait" without incurring a giant context-switch cost (SMT or Tera-style or Padauk FPPA-style; do GPUs do this too, at larger-than-warp granularity?). Event loops are a variant on this, and io_uring seems to think that's the future. But the GreenArrays approach is to decide that the limiting resource is nanojoules dissipated, not transistors, so just idle some transistors in a synchronous wait. Not sure if that'll ever go mainstream, but it'd fit well with the trend to greater heterogeneity.
You can have IPC that's faster than a function call if it's between cores.
Before we learned how to make them fast, perhaps. They do now tend to be very fast[0][1].
>One of the reasons for this could be high context switch latency.
As multiserver systems pass a lot of messages around, the important metric is IPC cost. Liedtke demonstrated microkernels do not have to be slow, with L3 and later L4. Liedtke's findings have endured fairly well[2] through time. It helps to know that seL4[3] has an order of magnitude faster IPC relative to Linux.
You'd need it to do a lot (think thousands of times) more IPC for the aggregated IPC to be slower than Linux.
>So I was thinking what could be done here.
I don't have a link at hand, but there's some involvement and synergy between seL4 team and RISC-V. I am hopeful it is enough to prevent the bad scenario where RISC-V is overoptimized for the now obsolete UNIX design, and a bad fit to contemporary OS architectures.
0. https://blog.darknedgy.net/technology/2016/01/01/0/
1. https://news.ycombinator.com/item?id=10824382
2. https://sigops.org/s/conferences/sosp/2013/papers/p133-elphi...
Citation needed. What kind of hit are we talking about? 5%? 90%? We have supercomputers from the future that have capacity to spare. I would be willing to take an enormous performance hit for better security guarantees on essential infrastructure (routers, firewalls, file servers, electrical grid, etc).
I am wondering if the performance will pan out in practice, as it doesn't seem to have a very deep pipeline, so getting high clockspeeds may be a challenge. In particular the 5 clock branch mispredict penalty suggest the pipeline design is fairly simple. Production CPUs live and die by the gate depth and hit/miss latency of caches and predictors. A longer pipeline is the typical answer to gate delay issues. Cache design (and register file design!) is also super subtle; L1 is extremely important.
I haven't published my latest work (end of the week) I have a minor bump to ~6.5 DMips/MHz - Dhrystone isn't everything but it's still proving a useful tool to tweak the architecture (which is what's going on now)
It seems that despite a lot of valid criticism against (System)Verilog, nothing really seems to be a on trajectory to replace it today. I'm not sure if that's purely inertia (existing tooling, workflows, methodologies), other HDLs not being attractive enough, or maybe Verilog is just good enough?
Inertia in tooling is a REALLY BIG deal - if you can't run your design through simulation, (and FPGA simulation), synthesis, layout/etc you'll never build a chip - it can take a 5-10 years for a new language feature to become ubiquitous enough so that you can depend on it en ough to use it in a design (I've been struggling with this using System Verilog interfaces this month).
If you look closely at VRoom! you'll see I'm stepping beyond some Verilog limitations by adding tiny programs that generate bespoke bits of Verilog as part of the build process - this stops me from fat fingering some bit in a giant encoder but also helps me make things that SV doesn't do so well (big 1-hot muxes, priority schedulers etc)
* From a high level what does your dev iteration look like?
* Getting instruction traces, timing and resimulating those traces
* Power analysis, timing analysis (do you do this as part of performance simulation) ?
* Do you benchmark the whole chip or specific sub units?
* How do you choose what to focus on in terms of performance enhancements?
* What areas are you focusing on now?
* What tools would make this easier?I currently run low level simulations in Verilator where I can easily take large internal architectural trace, and bigger stuff on AWS (where that sort of trace is much much harder)
I haven't got to the power analysis stage - that will need to wait until we decide to build a real chip - timing will depend on final tools if we get to build something real, currently it's building on Vivado for the FPGA target.
Mostly I'm doing whole chip tests - getting everything to work well together is sort of the area I'm focusing on at the moment (correctness was the previous goal - being together enough to boot linux), the past 3 months I've brought the performance up b y a factor of 4 - the trace cache might get me 2x more if I'm lucky.
I spend a lot of time looking at low level performance, at some level I want to get the IPC (instructions per clock) of the main pipe as high as I can so I stare at the spots where that doesn't happen
I'm using open source tools (thanks everyone!)
Would my fhourstones [1] [2] benchmark be of any use?
Are there GPL'd designs for PCIe, USB, etc, that could be used to incorporate this into a SoC design? If not, how much work is that compared to this?
Also, what other kind of technical considerations would be involved to make this into a "real" chip on something like 28nm?
So far I haven't needed USB/ether/PCIe/etc I've sort of sketched out a place for those to live - I think that for a high end system like this one you can't just plug something in - real performance needs some consideration of how:
- cache coherency works - VM and virtual memory works (essentially page tables for IO devices) - PMAP protections from I/O space (so that devices can't bypass the CPU PMAPs that are used to man age secure enclaves in machine mode)
So in general I'm after something uniquer, or at least slightly bespoke.
I also think there's a bit of a grand convergence going on in this area around serdes's which are sort of becoming a new generic interface PCIe, high speed ether, new USBs, disk drivers etc are all essentially bunches of serdes with different protocol engines behind them - a smart SoC is going to split things this way for maximum flexibility
So GPL'd IO blocks - This is a great question, and something I have definitely been asking myself! One thing to keep in mind is that IO interfaces like PCIe, USB, and whatnot have a Physical interface ("Phy" for short.) Those contain quite a bit of analog circuitry, which is tied to the transistor architecture that's used for the design.
That being said, A lot of interfaces that aren't DRAM protocols use what's known as a SerDes Phy (short for Serializer De-serializer Physical interface.) More or less, they have an analog front end and a digital back end, and that digital back end that connects to everything else is somewhat standardized way. So it wouldn't be unreasonable to try to build something like an open PCIe controller that only has the Transaction Layer and Data Link Layer. While there are various timing concerns/constraints when not including a Phy layer (lowest layer,) I don't think it's impossible.
The other big challenge is that anyone wanting to use an open source design will definitely want the test benches and test cases included in the repo (you can think of them like unit tests.) Unfortunately, most of the software to actually compile and run those simulations is cost prohibitive for an individual, because it's licensed software. Also, the companies that develop this software make a ton of money selling things like USB and PCIe controllers, so I'll let you draw your own conclusions about the incentives of that industry.
Even if you were able to get your hands on the software, the simulations are very computationally intensive, and contribution by individuals would be challenging ...though not impossible!
Despite those barriers, it's a direction that I desperately want to see the industry move towards, and I think it's becoming more and more inevitable as companies like Google get involved with hardware, and try to make the ecosystem more open. Chiplet architectures are also all the rage these days, so it would be less of a risk for a company to attempt to use an open source design.
I'd really be curious to hear Paul Campbell's take on this question though. He definitely knows a lot more than I do!
Here's a SerDes from Purdue. I don't think this particular design has been validated in silicon yet though.
Ibex needed to add a pass with sv2v https://github.com/lowRISC/ibex/tree/master/syn
It does have 'sticky' state bits for FP and I can see how I'll implement them - the big problem is not setting them (because they can be accumulated in any order as instructions hit the commit stages), it's how you test them that effectively becomes a synchronising point in a pipe where you spend all your time trying not to do that - everything has to stop and line up in order before you can sense that state reliably.
Normally we retire up to 8 instructions per clock - merging 8 bits of sticky state is some simple logic (some or gates, it doesn't matter what order they get processed in), merging 8 saved PCs (which one do you save?) is harder, it probably means a priority encoder (order probably does matter) and a 63 bit mux (still doable in a clock though)
In general though MIPS and RISCV are similar sorts of RISC architecture, they make some of the same design trade offs (no condition codes for example), RISCV avoids some of the mistakes (delay slots for example) - I'd guess they're about the same amount of work. I can imaging making a version of my CPU by switching out the instruction decoders (probably not really that simple).
As far as SoC it probably doesn't matter - that's more of an issue of which internal buses you choose to use for memory and peripherals
SoC stuff is probably unrelated to arcitecture -
1 - architectural - RISCV has a nice clean ISA, it's adding instructions quickly, CMOV is contentious issue there - I'm not an expert on the history so I'll let others relitigate it - it's easy to add new instructions to a RISCV machine, unlike Intel/ARM it's even encouraged - however adding a new instruction to ALL machines is more difficult and may take many years. But unlike Intel/ARM there IS a process to adopt new instructions that doesn't involve just springing them on your customers
2 - remember RISCV is a no-condition code architecture - that would make CMOV require 3 register file ports (the only such instruction that also requires an adder [for the compare]) - register file ports are extremely expensive, especially for just 1 instruction
3 - micro-architectural - on simple pipes CMOV is pretty simple (you just inhibit register write, plus do something special with register bypass) I'd have to think very hard about how to do it on something like VRoom! with out of order, speculative, register renaming - I can see a naive way to do it, but ideally there should be a way to nullify such an instruction early in the pipe which would mean some sort of renaming-on-the-fly hack
In many ways I guess it's a bit like a conditional branch that needs a write port - in RISCV, without condition codes, your conditional call relative branch distance will be smaller because the instruction encoding will need to encode 2-3 registers
Though, maybe instead of a conditional call, a conditional signal could do, which'd clearly give no expectation of performance if it's hit, simplifying the hardware effort required.
https://youtu.be/nlu0foF3WBk?t=182
I know, I'm leaning hard on that second "A" there. :D
What your view on that kind of design for a laptop/phone/tablet processor?
For reference, Digikey unit price for a VU440 floats around the $40-60k range
My next set of work in this area will be integrating an L0 trace cache into the existing BTC - that will help me greatly up the per-clock issue rate
https://moonbaseotago.github.io/talk/index.html
It would be nice to get actual performance numbers rather than just frequency scaled Dhrystone but I suppose we have to be patient.
Having said that I'm about reaching the end of the point where it's the only thing - being able to run bigger longer benchmarks is one of the reasons for bringing up linux on the big FPGA on AWS
I haven't actually looked at the generated code, but I imagine it's thousands upon thousands of instructions in a row with no conditional branching at all. CHOMP.
Be aware though, the micro-architecture used here is very interesting but differs in many ways from state of the art industrial high-end micro-architectures for superscalar out-of-order speculative processor.
I am quite curious about how the author came up with these choices
Seriously though I started out with the intent of building a 4/8 instruction/clock decoder, and an O-O execution pipe that could keep up - with the end goal of at least 4+ instruction s/clock average (we peak now at 8) - the renamer, dual register file, and commitQ are the core of what's probably different here
This looks like a renaming scheme used in some old micro-architecture (Intel Core 2 maybe) where ROB receives transient results and acts as a physical regfile, at commit reg value are copied to a arch regfile. But in your uarch the physical regfile is decoupled from ROB, which must correspond to your commitQ.
I wonder if this solution is viable for a very large uarch (8 way) because read ports to copy reg value from pysical regfile to arch regfile are additional read ports that can be avoided with other (more complex) renaming scheme. These additional read ports can be expensive on a regfile that already has a bunch of ports.
Any thoughts about this?
But I haven't read much of your code yet, that's just a raw observation
It does mean lots of register read ports .... but you can duplicate register files at some point (reducing read ports but keeping the write ports) (you want to keep them close to the ALUs/multipliers/etc) - in some ways these are more implementation issues rather than 'architectural'
Thank you for your opinion and thought process, it's very valuable !
BTW once great thing that sort of falls out of this architecture is that the commit register file gets shared between the integer and FP registers (and probably vector registers too), and duping just that may be an interesting architectural way to go
I understand how it applies to the HDL, but I doubt that it obligates you have to open your code to users of physical chips.
But I'm looking to find someone to build this thing, it's been a while since I last built chips (last CPU I helped design never saw the light of day due to reason that had little to do with how well it worked). So I need a way to show it off, show it's real. So GPLing it is a great way to do that - as is showing up on HN (thanks to whoever posted this :-).
In practice the RTL level design of a processor is only a part of making a real processor - a real VRoom! would likely have hand built ALUs, shifters, caches, register files etc those things are all in the RTL at a high level but are really different IP - likely they'd be entangled with GPL and a manufacturer might feel that to be an issue.
However I'm happy to dual license (I want to get it built, and maybe get paid to do it).
Also about half the companies building RISCVs are in China (I've been building open source hardware in China for a decade or so now, so I know there's lots of smart people there) - they have a real problem (in the West) building something like this - all the rumors about supply chain/etc stuff - having an open sourced GPL'd reference that's cycle accurate is a way help build confidence.
I spent a few years working on an x86 clone, I had maybe 10 (now expired) patents on how to get around stupidly obvious things that Intel had patented - (or around ways to get around ways to get around In tel that other's had patented) - frankly from a technical POV it was all a lot of BS, including my patents
This is a great strategy!
> I spent a few years working on an x86 clone, I had maybe 10 (now expired) patents on how to get around stupidly obvious things that Intel had patented - (or around ways to get around ways to get around In tel that other's had patented) - frankly from a technical POV it was all a lot of BS, including my patents
It might be worthwhile to GPL implementations of those expired patents if they are at all likely to be useful. And perhaps then do a bit of procedural generation of various combinations of them for release under the GPL as well (because those would be newly patentable).
Prior art FTW!
Probably not. It's always something stupid that can be easily worked around.
The real problem is convincing a jury: no one wants to risk hundreds of millions of dollars based on what 12 random people think. nVidia caved to Intel after building a Transmeta style VLIW chip that could run x86 assembly because a decade long patent battle would have been costly and invalidated patents on both sides.
I don't think the GPL anti-tivoization clause has much bearing there other than presumably you'd have to provide the full tool chain that resulted in the final gates - presumably this would affect companies producing actual chips the most since you couldn't have any propriety optimization or layout steps in producing the actual chip design, though also no DRM for FPGAs (is that even a thing?)
I think that’s the spirit of the GPL in a hardware context, but I don’t think it’s a given (by a long stretch) that courts would accept that argument.
A somewhat clearer case would be if you bought a device that implements a GPL licensed design in a FPGA. I think you could argue such devices cannot disable the reprogrammability of the FPGA.
If somehow this code is not in a firmware... No idea.
some of this design IS firmware - the lowest level bootstrap is encoded into an internal ROM - currently it's a very dumb bootstrap, a real implementation would boot from a number of possible sources. The sources are there on github.
All the ARM systems you can buy today have a similar embedded boot loader - almost all of them do not release that source, because it's the root of their secure boot chain.
IMHO this code should be public (but not the keys)
RMS wrote "I've considered selling exceptions acceptable since the 1990s, and on occasion I've suggested it to companies. Sometimes this approach has made it possible for important programs to become free software."
Computer Organization and Design, by the same authors, is considered a better choice for a first book. I personally loved it and couldn't put it down the first time I read it.
https://www.elsevier.com/books/computer-organization-and-des...
https://www.google.com/books/edition/Digital_Design_and_Comp...
This book definitely skews pragmatic, hands on and doesn't assume much. Covers both VHDL and Verilog. Has sections on branch prediction, register renaming, etc.
For that I highly recommend: https://www.cambridge.org/us/academic/subjects/engineering/c...
Great first book on the subject.
So much for "we will do only simplest of commands and u-op fusing will fix performance".
It is why I'm very suspicious about this argument from RISC-V proponents.
And such "conventions" are bad idea, like comments in code, IMHO. It can not be checked by tools, etc.
For some definitions of decent, I think that ship has sailed.
https://clang.llvm.org/docs/CrossCompilation.html
-target <triple> The triple has the general format <arch><sub>-<vendor>-<sys>-<abi>, where: arch = x86_64, i386, arm, thumb, mips, etc. sub = for ex. on ARM: v5, v6m, v7a, v7m, etc. vendor = pc, apple, nvidia, ibm, etc. sys = none, linux, win32, darwin, cuda, etc. abi = eabi, gnu, android, macho, elf, etc.
Note, none of those are exhaustive...
So, only "sub" is somewhat relevant and it is exactly what RISC-V should avoid, IMHO, and it doesn't with its reliance on things like u-op fusion (and not ISA itself) to achieve high-performance.
For example, performance on modern x86_64 doesn't gain a lot if code is compiled for "-march=skylake" instead of "-march=generic" (I remember times, when re-compiling for "i686" instead of "i386" had provided +10% of performance!).
If RISC-V performance is based on u-op fusing (and it is what RISC-V proponents says every time when RISC-V ISA is criticized for performance bottlenecks, like absence of conditional move or integer overflow detection) we will have situation, when "sub" becomes very important again.
It is Ok for embedded use of CPU, as embedded CPU and firmware are tightly-coupled anyway, but it is very bad for generic usage CPU. Which "sub" should be used by Debian build cluster? And why?
Edit: for grammar & typos
The extremum for this is getting a 10x performance boost by using, e.g., POPCNT, and suffering instead a 10-100x pessimization because POPCNT is trapped and emulated.
Are you guessing that the extension is optional specifically so that nobody will need to emulate things they can't afford to implement in hardware?
But trapping and emulating is explicitly allowed. Maybe it should be possible to ask at runtime whether an extension is emulated. Maybe it is? But I have not seen any way to tell. I guess a program could run it a thousand times and see how long it takes... It would be a serious nuisance to need to do that for each optimization, and then provide alternate implementations of algorithms that don't depend on the missing features.
This is why leaving popcount out of the core instruction set is such a nuisance. It is cheap in hardware, and very slow to emulate.
On organically-evolved ISAs, there are about N variants that correspond to releases. You can decide what is the oldest variant M you want to support, and use anything that is implemented in targets >=M; and the number of <M machines declines more or less exponentially with time.
With RISC-V, there are instead N=2^V variants, at all times, increasing exponentially with time. You too frequently don't know if your program might need to run on one that lacks feature X. So you (1) arbitrarily fail on an unknown fraction of targets, (2) fail on some and run badly on some others, with instructions you relied on for optimization instead emulated very slowly, (3) run non-optimally on all targets, or (4) have variant versions of (parts of) your program configured to substitute at runtime, for each of K features that might be missing. None of these choices is tenable.
The notion of "profiles" appears meant to reduce the load of this problem, but that makes it even more complicated.