RISC in 2022
wiki.alopex.li
wiki.alopex.li
I think the design principles of RISC were actually on a meta level above this: take the quantitative approach, use your transistors to best serve the software you have and compilers you can build. "Nice ISA to write assembly code for" was thrown out, or at least demoted significantly.
In that 80s moment in the transistor count curve, it meant simplifying the ISA very radically in order to implement fewer instructions in hard coded way without microcode, to the point of ditching HW multiply instructions. The microarchitecture could either do fast, pipelined execution or large instruction set in the transitor budget, you optimized for bang for the buck in the whole-system sense. You could make simple fast machines that were designed to run Unix (so had just enough VM, exception, etc support).
1. As microprocessors became dominant freezing complex microcode - with possible bugs - on an IC was a really bad idea. Better to run with simpler instructions and less microcode.
Look at the debugging issues that National Semiconductor had with the NS32016 - which I think really hindered its adoption.
2. You needed a much smaller team to design a RISC CPU - low double figures for IBM 801, Berkeley RISC, MIPS and Arm. This opened the door to lots of experiments and business models that would not have been possible with CISC.
What RISC took advantage of was decoupling the memory bus from the CPU clock rate and the introduction of instruction and data caches.
RISC is a simple interface, but the internal processor complexity is as high or higher than CISC (thanks to RISC simplicity). Speculative execution, reordering instructions, hazards, branch prediction these issues are the same across both.
The front end isn't this huge deal. CPUs all do the same thing, attempt to unroll huge state machines and compress time. The ISA is just a way to get that problem into the CPU.
I actually cited an example of a major CISC design that basically failed because it was so buggy.
The early microprocessors all had hard coded microcode - Intel only had upgradeable microcode with P6 in the mid 1990s.
I didn't say that it was the reason - there were many reasons - but that it was one factor. It was absolutely the case that original RISC designs were simpler to design and that helped them to get traction.
Edit - just to add that caches enabled RISC to get decent performance but they were not by themselves a reason for choosing RISC over CISC.
1990? No system without soft microcode? That's a very restricted definition of system for the field of computing in 1990.
Intel 8086 was introduced in 1978, with hardcoded microcode. Intel didn’t introduce microcode patching support until the Pentium Pro in 1995.
Interesting. Was the guiding principle of CISC ISAs essentially that be easy to write assembly code for then? Maybe it's obvious but I had never considered how or where CISC evolved from. Would I have to look at something like the history of the VAX ISA to understand this better?
What is certainly true is that a single instruction could do a lot, making for much more concise code than would be the case for RISC code.
IBM S/360 is probably the most influential CISC architecture. There is lots of S/360 documentation online. If of interest I did a short post on S/360 assembly a little while ago.
https://thechipletter.substack.com/p/writing-ibm-s360-assemb...
I genuinely had no idea how awful it was - of course you don't see pop-ups if you are logged in. To be fair to Substack the control to turn off pop-ups was there but it was turned on by default and a bit buried away in the UI.
I've removed any mid-article calls to action and turned off pop-ups now so hopefully a much better and less aggressive experience.
On the S/360, the article I linked to was a short look at how complex some S/360 instructions were and at how assembly used to be written using pen and paper. If you're interested in more on S/360 Assembly then the principles of operation may be worth a scroll through. [1]
Thanks again - really appreciate that you fed back rather than just closing your browser window.
[1] http://bitsavers.org/pdf/ibm/360/princOps/A22-6821-0_360Prin...
And more importantly, future compilers, not the tech they had at the time.
You have it right: this is almost exactly (in different words!) what Radin wrote in his original RISC paper.
The article doesn't get RISC-V instruction encoding right. It mentions compressed instructions, but instructions can also be longer than 32 bits. The important thing about RISC-V is that the instruction stream can easily be divided at instruction boundaries (unlike, say, x86 which is horrific to decode). This gives you most of the benefits of fixed size instructions and the benefits of extensibility when you need it.
Who'd have thought thirty year ago we'd all be sittin' here talking tens of registers, eh?
In them days we was glad to have two or three.
If we were lucky.
My first computer had 1 (one). One with less bits than fingers on our hands. It was so dear to us we even gave it a name. "Accu" it was called.
If you tell that to the young people today, they won't believe you.
The entire article is about instruction sets, not about their physical implementation, so I don't think this clarification is necessary.
In a modern high performance processor instructions are decoded in batches: Decoding the first instruction is straightforward. But x86 instructions range from 1 to 15 bytes, therefore the second instruction can start from byte-offset 1 up to 15. 3rd instruction has a byte-offset ranging from 2 to 30, ans so on. Furthermore, figuring out an x86 instruction length requires reading several byte from the instruction.
In the end, the 8th instruction has 99 possible byte-offset, and assuming that we put, as you suggest, a decoder for each position and length, we need about 1590 decoders and many multiplexer to decode 8 full instructions per cycle.
Of course we don't do that, it would consume a lot of energy for nothing.
To handle that, modern x86 processor instruction decoding involves a instruction length decode before the instruction decode. The instruction length decode is responsible for identifying the instruction positions and boundaries, and this instruction length decode is a challenging part of the x86 processor to design. We don't know how Intel or AMD exactly do instruction length decode, but we know that some published techniques include a length predictor.
That's why, for simplicity and energy efficiency, instruction boundaries must be easily identified and the number of instruction lengths must be kept low.
Um... wat? No CPU tries to decode 99 bytes of memory in a cycle. ADL is at 32 currently, I believe. And the instruction starting at byte 12 doesn't change depending on anything but it's own data. It either exists (because the previous instruction ended on byte 11) or it doesn't. So you decode 32 instructions starting at each byte you've fetched (the last ones can be smaller subset engines because they don't need to decode longer instruction forms), and then mask them on or off based on earlier instruction state. Then feed your 1-32 decoded instructions through a mux tree to pack them and you're done.
Surely there's more complexity, since this is going to have to be pipelined in practice, and a depth of 32 is going to require something akin to a carry-lookahead adder instead of being chained.
But the combinatorics you're citing seem ridiculous, I don't understand that at all.
Actually, no x86 processor decodes 8 instructions in parallel. This is an example to illustrate how the number of possible offsets scales with 15 instruction lengths.
> So you decode 32 instructions starting at each byte you've fetched
No you don't do that, it's too power consuming.
> But the combinatorics you're citing seem ridiculous, I don't understand that at all.
What I'm trying to explain is that decoding 8 instructions in parallel in x86 is hardly possible, while decoding 8 instructions (or more) from a RISC archi per cycle is never a problem
Uh... yes you do? How else do you think it works? I'm not saying there's no opportunity for optimization (e.g. you only do this for main memory fetches and not uOp execution, pipeline it such that the full decode only happens a stage after length decisions, etc...), I'm saying that it isn't remotely an intractable power problem. Just draw it out: check the gates required for a 64->128 Dadda multiplier or 256 bit SIMD operation and compare with what you'd need here. It's noise.
And your citation of "8 instructions in parallel" seems suspicious. Did I just get trolled into a Apple vs. x86 flame war?
No, I literally explain it in my first answer. The part about "1590 decoders" is irrelevant since a misunderstood your message (thinking that you are talking about using 16 decoders to decode the 16 instruction lengths of a single instruction).
But the rest on instruction length decode is how you actually do it.
> I'm saying that it isn't remotely an intractable power problem.
I mean, obviously, if you ignore all the power consumption issues of using 32 decoders in parallel and using only 5 of the results out of the 32. Then yes, there's no problem.
But in reality, yes it's a problem to decode many x86 instructions in parallel.
> Just draw it out: check the gates required for a 64->128 Dadda multiplier or 256 bit SIMD operation and compare with what you'd need here. It's noise.
Yes, the energy consumption of the multipliers is high, but I don't see how this is an argument to make an inefficient decoder? Also, a multiplier power consumption depends on transistor activity, and you can expect the MSB of the operand not to change too much. For decoder the transistor activity will be high.
> And your citation of "8 instructions in parallel" seems suspicious. Did I just get trolled into a Apple vs. x86 flame war?
Not a troll nor a flame war. I don't use Apple products, mainly because I don't agree with Apple practices. But actually choosing a RISC ISA allows them to decode a lot of instructions in parallel for little energy and complexity.
I chose 8 because it is the maximum that the mainstream will currently see. You might argue that 8 RISC instructions are not comparable with 8 CISC instructions, but even with say 4 CISC instructions it will still consume more energy
Alder Lake decodes six. And again, your intuition about power costs here is just simply wrong. Instruction decode is Simply Not a major part of the power budget of a modern x86 CPU. It's not.
I never said that instruction decode was a major part of the power budget.
And precisely, it is not because they don't decode 32 instructions in parallel. That's the purpose of an instruction length decoder prior to instruction decode.
> RISC was a set of design principles developed in the 1980’s that enabled hardware to get much faster and more efficient.
There is a strong argument that RISC as a set of design principles (if not as an acronym) started in the 1970s with the IBM 801 [1] and many of the ideas date back to the 1960s with the CDC 6600 mainframes which were very RISCy.
On a more substantive point, I don't think small code size was ever 'officially' part of the RISC concept. Arm pioneered it with Thumb but I think that was a pragmatic decision to get Arm into devices with limited memory space such as early mobile phones.
[1] https://thechipletter.substack.com/p/the-first-risc-john-coc...
This does not seem to hold for SPARC: according to
> https://www.cl.cam.ac.uk/~pes20/weakmemory/x86tso-paper.tpho...
the (strong) memory models of x86 and SPARCv8 are very related:
"We give two equivalent definitions of x86-TSO: an intuitive operational model based on local write buffers, and an axiomatic total store ordering model, similar to that of the SPARCv8."
"Our x86-TSO axiomatic memory model is based on the SPARCv8 memory model specification [20, 21], but adapted to x86 and in the same terms as our ear- lier x86-CC model."
"We have described x86-TSO, a memory model for x86 processors that does not suffer from the ambiguities, weaknesses, or unsoundnesses of earlier models. Its abstract-machine definition should be intuitive for programmers, and its equiva- lent axiomatic definition supports the memevents exhaustive search and permits an easy comparison with related models; the similarity with SPARCv8 suggests x86-TSO is strong enough to program above."
Actually only 18. With 21 bits you can get 128 registers.
Another divot: asymmetric functional units. Some versions of Alpha supported a PopCount instruction, but it only worked in a single functional unit, which made scheduling a pain, esp. if you had to write in assembly language.
I'm not convinced that AVX 256 and AVX 512 are useful for non-matrix operations. Most strings (more importantly, parsing bounded by whitespace) are much shorter than 512 bits (32 bytes). In English, I cannot come up with many words longer than 16 bytes (some place names, antidisestablishmentarianism, chemical compound names, and some other stuff)
I've observed that compared to regular x86-64 code without SIMD, using AVX 256 speeds up the Chacha20 cipher (for long messages so they can be processed in 512-bytes chuncks (8 blocks)) by a factor of 5. Network packets easily exceed 1KB, and files are usually much bigger.
Matrix operations aren't the only viable niche.
https://www2.eecs.berkeley.edu/Pubs/TechRpts/2016/EECS-2016-...
One of my personal favourite quotes comes from the foreword to the PA-RISC v2.0 manual[0] by Michael Mahon, the PA-RISC v2.0 principal architect:
> Efficiency also has evident value to users, but there is no simple recipe for achieving it. Optimizing architectural efficiency is a complex search in a multidimensional space, involving disciplines ranging from device physics and circuit design at the lower levels of abstraction, to compiler optimizations and application structure at the upper levels.
> Because of the inherent complexity of the problem, the design of processor architecture is an iterative, heuristic process which depends upon methodical comparison of alternatives (“hill climbing”) and upon creative flashes of insight (“peak jumping”), guided by engineering judgement and good taste.
> To design an efficient processor architecture, then, one needs excellent tools and measurements for accurate comparisons when “hill climbing,” and the most creative and experienced designers for superior “peak jumping.”
Engineering and good taste! – we do not come across those very often.
Aarch64 also has a dedicated SP register.
I don't think this is correct. It has a suggested SP register, which merely gets some special support in the C subset. But that's just an optional compression scheme and not really part of the ISA design.
As far as I know, that's the full extent to which RISC-V has a dedicated stack register: it has a compressed instruction format that uses x2 as a base register, but not in the base ISA, just a standard extension. There's no dedicated PUSH or POP instruction, no dedicated instruction for storing the link register into the stack, no dedicated instructions for incrementing or decrementing x2 (you do that with ADDI, which can be compressed as C.ADDI as long as the stack frame size is less than 32 bytes, which means it has to be 16 bytes in the standard ABI), not even autoincrement and autodecrement addressing modes.
Basically there are special registers for trap handling, which are CSRs: xscratch (a scratch register), xepc (the trapping program counter), xcause and xtval (which trap), and xip (interrupts pending). These come in four sets: x=s (supervisor-mode), x=m (machine-mode, with a couple of extras), x=h (hypervisor mode, which has some differences), and x=vs (virtual supervisor). You can't handle traps in U-mode, so in a RISC-V processor with trap handling and without multiple modes, you're always in M-mode. (See p.3, 17/155.)
I haven't done this but I suppose that what you're supposed to do in a mode-X trap handler is start by saving some user register to xscratch, then load a useful pointer value into that user register off which you can index to save the remaining user registers to memory.
I guess you know xscratch (and xepc, etc.) wasn't previously being used because you only use them during this very brief time and leave x-mode traps disabled until you finish using it. If all your traps are "vertical" (from a less-privileged mode like U-mode into a more-privileged mode like M-mode) you don't have to worry about this, because you'll never have another x-mode trap while running your x-mode trap handler.
I should probably check out how FreeBSD and Linux handle system calls on RV64.
dh` explained the following technique to me, as explained to him by jrtc27: upon entry to, say, an S-mode trap handler, you use CSRRW to swap the stack pointer in x2 with the sscratch register, if it's null you swap back, then push all the registers on the stack, then you can do real work.
BTW. Some MIPS processors did have a FMA instruction that did round twice. The compiler was thus able to fuse instructions without the code giving different results. This was deprecated in later versions, however.
Except when it's time to boast about FLOPS. Then it's always counted as two operations. :-)
I wonder, if it would make sense to decouple compression and microcodes. So you could take a body of Code and find the best compression for it. Or even be able to change "lookup tables" before starting the operating system. Possibly have different compression methods (x86 / arm..) run on the same CPU, without any drawbacks. Which could get you around licensing an ISA. (Yes I'm a software engineer thinking about hardware)
The “without any drawbacks” part never panned out, though.
It also is fairly similar to what Apple has done with emulating 68k on PPC, PPC on x64, and x64 on ARM (but those, AFAIK, do not offer full emulation of the host CPU, as they don’t need to run code in kernel mode)
This looks to have started when FPUs were optional and/or physically separate, but that's no longer the case.
Any examples-of/thoughts-on of systems that normally put floats and integers in the same registers?
This makes MMX very unpopular outside controlled situations, so compiler autovectorization doesn't support it.
Using a separate register set for FP is not just about making floats optional. It also allows to better isolate the float and int units and to build a more efficient micro-architecture.
For example: using a single physical register bank for floats and integers would be expensive (as the size of the register bank grows quadratically with the number of read/write ports), therefore using separate physical register bank for float and integer is more efficient.