How to Design an ISA
queue.acm.org
queue.acm.org
L1i matters, people!
RISC-V consistently wins on L1i footprint.
The complaining is about number of dynamic instructions ("path length"), which can hit you if you don't fuse. Of course, path length might not actually be the bottleneck to raw performance, but it's an easy metric to argue, so a lot of people latch on to it.
Ironically, RISC-V does great there[0]. Note this is despite these researchers did not even consider fusion.
That's awesome.
ARM of course would also benefit from fusion too; but camel-cdr's mention of it being only rv64g is a pretty significant caveat.
No, winning 4 and losing 6, by a small margin, isn't "being worse than arm". The paper's authors even explicitly conclude it is not losing to ARM.
This is even ignoring whether code is within or outside loops, counting fuseable instructions as always non-fused, and not considering any instructions from extensions after 2019's ratified (actually unchanged from 2017) rv64g... any of those would have a favorable effect on RISC-V.
This is an excellent result for RISC-V, that clears any doubts in terms of path length. On top of what we already know about RISC-V leading in code density in 64bit.
Excluding extensions is perhaps a significant question, but, for example, Debian RISC-V currently targets rv64gc, which should have the same instruction counts as rv64g does, so software compiled for Debian can't use the later extensions for most code anyway. (never mind that ARMv8 also has excluded extensions, namely NEON, which is always present on ARMv8 and is not designed to be ignored)
And, of course, even being better than ARM is not equivalent to being the best it could be; ARMv8 isn't some attempt at a magical optimal instruction set, it's designed for whatever ARM needed, and that includes being able to efficiently share hardware with ARMv7 for backwards compatibility.
Simplicity has enormous value.
> When it started testing simulations of early Pentium prototypes, Intel discovered that a lot of game designers had found that they could shave one instruction off a hot loop by relying on a bug in the flag-setting behavior of Intel's 486 microprocessor. This bug had to be made part of the architecture: If the Pentium didn't run popular 486 games, customers would blame Intel, not the game authors.
Does anybody know the details of this?
> Sadly, my source for this was a former Intel chief architect, and I don’t think he ever said it anywhere that was recorded. […]
If you're more interested in just concrete examples of this kind of thing, apparently IBM's System/360 team ran into heaps of this kind of issue when emulating the IBM 1401. It's mentioned in Frederick P. Brooks, Jr.'s book The Mythical Man-Month in the Formal Definitions section of Chapter 6, and I think he probably discussed it in more detail in his (and Blaauw's) book Computer Architecture.
Edit: pasting in a cleaned up version of the relevant part of the transcript:
1:04:56 It's required to run every one of those well, how do I know it does? I can't test them all, so you say well you as long as you design to the architecture spec that should be enough. Right? Ha no for example, inside the architecture spec there are places where it says this condition flag is undefined as a result of this operation so you'll do – I don't know what it was anymore, add operation or something, no, it can't be add pick a different [thing] – there was some instruction that would say I do not guarantee what the carry flag will look like when I'm finished and you go as an architect hey that's cool it means I can do it either way whatever way is easiest. No it doesn't.
Yeah if you think that you're going to get in big trouble. Because what what will happen is – and this literally happened which is why I know about this – you put the chip out and then you discover oh it was easiest for my team to set the bit to a 1 didn't matter because it was undefined right I get to pick, but all the previous chips were setting it to a zero although they were calling it undefined. Now you're in trouble, because what you're going to discover some goofy app out there required that the bit be a zero after that operation even though the book said it was undefined. And they didn't notice because up until now it always was a zero. But your chip comes out, the software doesn't work anymore. Guess who's at fault? You are. Can you go "hey look at what the book says, can't you read?" and they'll say "I don't care what you say your chip doesn't work my software, you're a loser, your chips busted."
[1] https://youtu.be/jwzpk__O7uI?si=iy23ZM5tQX-hI87C&t=3903
[2] https://www.sigmicro.org/media/oralhistories/colwell.pdf
So what did you do w.r.t. addressing modes? Do you have memory-memory operations? Is it a load/store architecture?
I found the CPU specification (https://github.com/vircon32/Vircon32Documents/blob/main/Spec...), but I couldn't find an ISA specification.
There are no capabilities like DMA, so memory-memory operations are not really possible with the only exception of instruction MOVS which acts as a supposed MOV [DR], [SR].
Resumable instructions are clearly also a big deal.
https://people.eecs.berkeley.edu/~krste/papers/EECS-2016-1.p...
I actually started the design of the MRISC32 ISA before I knew about RISC-V (I even called it VRISC first, for "Vector-RISC", but had to give that name up for obvious reasons).
Initially I attacked the problem from a software developer perspective, trying to keep an open mind towards various possible solutions, but over the years I have learned lots of things and have scrapped many ideas that simply would not work well in hardware. Developing an FPGA implementation in parallel with developing the ISA certainly helped alot.
One of the luxuries of running it as a one-man open source project is that you are not bound by business deadlines and goals, so it's perfectly OK to change your mind half-way in, and let ideas and concepts mature as you learn more. And unlike the very large body that is RISC-V, you don't have to struggle with competing interests either.
I wonder if learning an optimal ISA from given ones would require optimizing 1. time semantics, 2. instruction semantics, 3. source code under uncertain distributions from given examples or if it is completely intractable.
But the big question I can't answer is: is this somehow a really bad ISA for high end CPU's? I think it's likely, as likely 30% more instructions are needed.
This is an intense 15 part series, from before the author became well known. It takes one through the SuperH ISA from the viewpoint of a compiler writer, for Microsoft Windows CE.
Even the SuperH folks seem to have realized this since later versions do use 32bit insns with 16bit as a special case, much like Thumb2 and RISC-V. Some ISAs have 24bit insns but these are rather clunky in other ways.
But sure, there is always demand for new instructions, and there is limited space. I imagine things like SIMD could be shoved into the secondary instruction set. Though I imagined that secondary ISA would by default be a Forth jump list, it would come virtually for free.
ISA principles are covered in Hennessy and Patterson's Computer Architecture: A Quantitative Approach. But then they've been relegated there out to an appendix. In addition to not being formally taught, they're de-emphasized.
> If you buy an NVIDIA GPU, you do not get a document explaining the instruction set. It, and many other parts of the architecture, are secret. If you want to write code for it and don't want to use NVIDIA's toolchain, you are expected to generate PTX, which is a somewhat portable intermediate language that the NVIDIA drivers can consume. This means that NVIDIA can completely change the instruction set between GPU revisions without breaking your code. In contrast, an x86 CPU is expected to run the original PC DOS (assuming it has BIOS emulation in the firmware) and every OS and every piece of user-space software released for PC platforms since 1978.
Maybe a few minor issues would end up corrected, or maybe China will fork the project to do so, but that's about it. Compatibility will reign, and RISCV is probably the last ISA you'll need to know. Thankfully.
First of all I don't think that "compatibility will reign". It's more like once the industry really starts picking up RISC-V, fragmentation will reign (at least for a decade or so).
It's also quite likely (IMO) that we'll see a "next generation" rather sooner than later, i.e. "RISC-VI". RISC-V, with some agreed upon extensions, may become the norm for Android, mobile, automotive and so on, but for the high end (servers, gaming, etc) I think that the industry will push for a different philosophy than the RISC-V authors originally envisioned - and that could become a new "revision" if you will.
There is very little room to complain about what RISC-V does have, in RV32I/RV64I or even RV32G/RV64G. The core ISA works just fine and its primary attribute is that everyone is legally free to use and build on it, and that there is a large and growing body of software that runs on it.
The complaints are about things it doesn't have. No carry bit. No complex addressing modes. That kind of thing. If someone proves that those actually matter, for example by building a CPU with custom instructions that blows away everyone else's, then those can be added to RISC-V. There should never be any reason to need a "RISC-VI".
I'm actually quite sympathetic to Qualcomm's proposal to add some of the things Aarch64 has, as an optional but standardised extension. On the other hand I'm completely against the second part of their proposal, to do a "big bang" replacement of the C extension with their extension in e.g. the RVA23 profile.
RISC-V cares about its trademarks; a non-compliant core would not be allowed to use them.
The rest is a software problem. I know that e.g. Linux exposes the required information about the CPU's ISA via a dedicated syscall.
I do not know whether there's some ELF header or the like.
Maybe it's out of selfishness (e.g. because they are repurposing an existing microarchitecture for RISC-V), or maybe it's because that's how they want to build their hardware (e.g. for their particular performance target decoding a plain-old 32-bit instruction may be more silicon/power efficient than to fuse 2-3 16-bit instructions).
Whatever the reasons, I think that RISC-V will have to live with this critique for as long as it lives.
I'm not saying that those are poor design choices, but they will always be pain points (for small cores and big cores alike - but for different reasons).
I don't think that you should underestimate the drive to modify an architecture if it does not fit your needs - especially if it is a free and open architecture like RISC-V. For example LoongArch has already happened, and I can easily see how something similar can happen if a major player decides to move from x86 or ARM to "something else" (e.g. if NVIDIA wants full control over their next gen super AI solution).
I don't think we'll see a repeat of Qualcomm's attempt. This was a very special situation that they ended up with the NUVIA purchase/ARM lawsuit fiasco.
Fortunately, RISC-V foundation handled the situation well. As RISC-V continues to grow exponentially, an unlikely later attempt will meet even stronger resistance, not just from the set precedent, but from the larger already deployed ecosystem of software and hardware.
It's manageable and you can live with it, but already from the start you have an unnecessary legacy that needs to be handled.
If there is enough consensus in the industry, a new revision may be the best way forward.
The comment (https://lobste.rs/s/v8xovv/how_design_isa#c_pluxuy) about JALR doesn't convince me. Yes you can use other registers for millicode and coroutines, but that is also baked into the ABI. See "Return-address stack prediction hints encoded in the register operands of a JALR instruction." in the unprivileged spec.
It's an argument for supporting 2 link registers, not 32. I suppose you could argue "but there might be a future use that needs 3!" but it's definitely not clear cut.
I haven't seen a solid defence for the lack of conditional move or advanced addressing modes either.
Also the inclusion of compressed instructions in RVA22/23 seems to be a mistake: https://lists.riscv.org/g/tech-profiles/topic/101741936#297
I don't think Qualcomm have published their proposal & benchmarks in full, but I can the pain of compressed instructions is very very very high so it would be worth removing them from RVA22/23 even if there is a slight performance penalty. Though it is probably too late realistically, since the profiles are meant to be backwards compatible.
Overall I still think these are pretty minor mistakes and RISC-V is a nice ISA.
That's backwards. It's the people who want those who need to prove they make a significant difference, and not just with hand-waving but with actual chips and data. "Everyone else does it" is not data.
Thus far, RISC-V cores come in very competitive and even faster than Arm cores with similar µarch e.g. SiFive U74 vs Arm A55, or THead C910 vs Arm A72.
> the pain of compressed instructions is very very very high
Only for people who bought a company with a fast Aarch64 core that Arm is suing them over using so are trying to convert it to be a RISC-V core instead, with minimal work.
The companies that are designing fast and wide RISC-V cores from scratch (Ventana, Tenstorrent, Rivos, ...) are saying the C extension is no big deal to implement, and of course does have real advantages in static and dynamic code size, icache size / performance / bandwidth etc.
The discussion you referenced has Qualcomm claiming Rivos is also against the C extension, and then a Rivos person coming back saying "Hang on just a minute ... we're fine with C, we're just open-minded enough to want to see real data on your proposal".
This has an extensions that support conditional move and indexed memory access :-D
See https://sourceware.org/binutils/docs/as/RISC_002dV_002dCusto...
In any case you aren't going to be able to see the benefit by comparing totally different cores. There are too many other factors. The article says:
> Arm considered eliminating predicated execution entirely, but conditional move and a few other conditional instructions provided such a large performance win that Arm kept them.
It would definitely be great to see numbers here but I see no reason to doubt that.
> Only for people who bought a company with a fast Aarch64 core that Arm is suing them over using so are trying to convert it to be a RISC-V core instead, with minimal work.
I'm not sure what you are talking about here. I was referring to the fact that the C extension means uncompressed instructions may not be naturally aligned which leads to all sorts of complexities, e.g. fetching instructions that are split over a page boundary, different PMA regions, different PMP regions, etc. It adds a lot of complexity to CHERI too. It's not a "big deal" to implement, but it does add significant complexity which would have been nice to avoid if it wasn't actually necessary.
RISC-V has Zicond now.
That has to go into my list of favorite quotes!
I. e. it's stupid and embarrassing to watch people use it.
You know, "I resemble that remark!" :-)
(But that doesn't change the fact that I actually like the Bjarne Stroustrup quote!) <g> :-) <g>)
Now let's understand your comment a little bit better. Your comment apparently arises from a Meme, specifically this one:
https://knowyourmeme.com/memes/we-should-improve-society-som...
While that meme is indeed entertaining(!) -- it is in no way actually relevant to the Bjarne Stroustrup quote!
It is a "Motte and Bailey" AKA, "bait-and-switch" argument/comparison.
Propagandistic agenda-driven AI chatbots seem to do this a lot -- but I'll be charitable (this time!) and assume, for the purposes of discussion, that you are human...
Perhaps we should all learn about what a "Motte and Bailey" argument/comparison/logical fallacy, is:
https://rationalwiki.org/wiki/Motte_and_bailey
>"Motte and bailey (MAB) is a combination of bait-and-switch and equivocation".
https://en.wikipedia.org/wiki/List_of_fallacies#:~:text=Equi...
https://pressbooks.ulib.csuohio.edu/eng-102/chapter/fallacie...
Phrased another way (in Billy Madison terms): "Mr. Madison... Everyone in this room is now dumber...":
https://www.imdb.com/title/tt0112508/characters/nm0235999#:~....
Incidentally, another one of my favorite Bjarne Stroustrup quotes is:
"Proof by analogy is fraud". :-)
https://www.stroustrup.com/quotes.html#:~:text=Proof%20by%20...
That's the prerequisite for me to engage any further with you.
Would RiskV be amenable to have language-specific customizations or peripheral accelerators?
Machine learning asics are prone to having instructions that correspond to activation functions. tanh, sigmoid etc.
A reasonable approach is to take a risc ISA and add domain specific instructions onto it. That gets you straightforward codegen and implementation for the 90% case and magic instructions to make the important path very fast.
Not wrong, but a lot easier in recent years as compared to the past, as historically there were a lot of closed-source systems. Nowadays, with the popularity and ubiquity of open source, one can 'brute force' writing for Linux, GCC, and LLVM, and you've probably a sizeable portion of use cases covered.
(You may have difficultly in getting things into mainline if you've got a niche ISA, so there's continuing overhead of maintaining patches.)
ISA's of old would have an "ADD" instruction that added two numbers.
New ISA's should have crazy complex instructions that implement things helpful to make javascript/python/whatever run fast.
We should set mostly-automated design-space-search programs off to consider millions of autogenerated new instructions and simultaneously figure out how that would impact compilers, CPU design, power consumption, performance, etc.
We are nowhere near that point though. Probably would only happen if we have a grand unification of CPU architectures.
IIRC, The M series processors were explicitly designed with extra instructions for commonly used JS functionality. Compilers for that platform can rely on such instructions for other languages as well.
Also, SPARC was designed to run C code fast - a function call would require little more than moving the register window to preserve the caller's context. Because of that, calling a function had a very small penalty compared to other architectures of the time.
I think I know what you're thinking of and it's a little overblown. It's a single instruction - FJCVTZS - which performs a specific type of floating-point to integer conversion. It's a standard ARMv8.3 instruction, not an Apple extension.
https://developer.arm.com/documentation/dui0801/l/A64-Floati...
https://github.com/YosysHQ/picorv32/blob/master/picorv32.v#L...
(this is the 'subtract' alu instruction decoder in pico riscv for example)
What if x86 CPUs werent designed for low power? e.g due to focus on competitivness in perf focused markets
Saying that one isa is faster or more energy efficient is like saying that c++ syntax is faster than java syntax.
While there are lang features that enable stuff, then almost everything is up to the implementation - compiler, libraries, runtime and the programs code.
One letters arent faster than the other. ISA doesnt imply perf. characteristics of the end product.
Read this:
https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-...
Even if you start talking about decoders, then:
>Another oft-repeated truism is that x86 has a significant ‘decode tax’ handicap. ARM uses fixed length instructions, while x86’s instructions vary in length. Because you have to determine the length of one instruction before knowing where the next begins, decoding x86 instructions in parallel is more difficult. This is a disadvantage for x86, yet it doesn’t really matter for high performance CPUs because in Jim Keller’s words:
Kind of.
ISAs don't exist in a vacuum - for a given transistor budget, they'll force chip design choices that will drive power consumption and performance. Decoding instructions is one thing, but reordering them quickly and efficiently is more impactful for both power (if it can be done with fewer transistors) and performance (if it can be done better/faster so that more instructions from more instruction flows can be retired at the same time).
I designed a beautiful ISA in college. It was a (mostly) stack machine with instructions designed to make a FORTH compiler extremely easy to implement. Unfortunately, it wouldn't be easy to evolve it past the point processors got faster than memory (I did not see that coming). It would, as originally designed, end up being unavoidably slow unless some fairly complicated caching were to be implemented.
Another interesting example is the Intel 432 and its bit-aligned instructions. A lot of silicon that could be better used elsewhere was dedicated to fetching instructions. It was also slower to implement.
On the x86 not being designed for efficiency, Intel has a whole line of CPUs designed for low-power environments. At some point, there was even a Motorola phone running Android on x86.
Note it failed in the market, and was never competitive.
Conversation went off the rails.
Which may be true, if we’re talking about compilation speed, although in this case the reverse is true.
I think that could be a valid statement. APIs can influence performance by constraining the implementation. For instance, the syntax for constructing an object in C++ will, generally speaking, always yield faster code than Java, because Java objects are almost always allocated on the heap, while C++ objects can be allocated on the stack. Compare:
// C++
MyObj o{};
// vs. Java
MyObj o = new MyObj();
Sure, it's possible to write a Java allocator/GC that will yield similar performance to the C++ code, but in general, that will practically never be the case. The syntax of the language has constrained the implementation so that Java will practically always be slower. Presumably, similar design choices in an ISA could have the same effect.Didnt you just agree with me that it is dependent on the impl/end product?
Because what would be the reasons in isa world to make it not desirable
>The syntax of the language has constrained the implementation so that Java will practically always be slower.
The most interesting question is: by how much?
1% 3%? 30?
No. I'm saying that it may be theoretically possible to tune performance in some cases, but not practical.