The RISC Wars Part 1: The Cambrian Explosion
thechipletter.substack.com
thechipletter.substack.com
CISC (heavily encoded) instruction sets made sense when memory was very expensive (late 70s I worked somewhere where we spent >$US1M for 1.5Mb of actual core for a mainframe - reading a word took 1uS - reading an instruction was destructive, you had to write it back) heavily encoded instructions made sense because of the realities of the memory hierarchy.
As we started to move CPUs onto single chips the numbers started to change, RAS/CAS timers were ~1-300nS, caches in particular, initially off chip and eventually on chip were in the 10ns sorts of speeds, suddenly the tradeoffs between heavily encoding instructions to make them small and trading off decode time for fetch time were reasonable changes to make
Interesting historical fact there is that early ARM cores (and I suspect that Berkeley RISC is similar) are what one would today call CISC microarchitecture with the datapath strikingly similar to lets say m68k, but with significantly simpler sequencing logic.
> Although the ARM2 employed by current models could reportedly be run at 20 MHz, it was only ever run at 8 MHz due to external limitations, these being the speed of the data bus and of the "relatively slow", but correspondingly relatively inexpensive, RAM devices in use. The ARM3 incorporated a 4 KB on-chip combined instruction and data cache, loosening such external constraints and thus permitting the processor to be run productively at the elevated 20 MHz frequency.
https://en.wikipedia.org/wiki/Acorn_Archimedes#ARM3_upgrades
Though it did support page mode DRAM.
> Another change, and among the most important in terms of practical real-world performance, was the modification of the instruction set to take advantage of page mode DRAM. Recently introduced, page mode allowed subsequent accesses of memory to run twice as fast if they were roughly in the same location, or "page", in the DRAM chip. Berkeley's design did not consider page mode and treated all memory equally. The ARM design added special vector-like memory access instructions, the "S-cycles", that could be used to fill or save multiple registers in a single page using page mode. This doubled memory performance when they could be used, and was especially important for graphics performance.
https://en.wikipedia.org/wiki/ARM_architecture_family#Design...
> ARM1 was distributed as an evaluation system and was never commercialized.
But well, many RISC designs had off-chip SRAM caches, sometimes even ridiculously large (PA-RISC is a prime example).
And?
You only NEED instruction fetches when the values you are operating on are in registers -- because you have a lot of registers, function arguments and results are passed in registers, local variables are in registers, even the return address stays in a register for leaf functions (the vast majority of functions, dynamically.
Loads and stores are quite rare, and the majority of them are saving registers at the start of the function and reloading them at the end -- for which ARM provided load/store multiple instructions, the "CICSy" part of ARM (but not very .. it's a simple hardware sequencer) which they have dropped in Aarch64.
I think we bought our first 'big' machine with semiconductor memory around then (a Vax 11/780)
The issue with this "debate" is that it misses the forest for trees. Instead we should be talking about binary encoding (ie how much "variability" is required), and you're right on that bit; memory isn't the issue it once was.
Citation needed. It's a load/store ISA, with arithmetic taking place only register-to-register.
That's assuming that CISC programs are smaller than RISC ones, which is not the case, at least for RISC ISAs with two instruction lengths (CDC6600, Cray 1, first IBM 801 version, RISC-II, and ARMv7 and RISC-V obviously)
As we started to move CPUs onto single chips the numbers started to change, RAS/CAS timers were ~1-300nS, caches in particular, initially off chip and eventually on chip were in the 10ns sorts of speeds, suddenly the tradeoffs between heavily encoding instructions to make them small and trading off decode time for fetch time were reasonable changes to make
...and now, CISC still makes sense because of the huge gap between core and memory speeds, accompanied by the many levels of caching in the middle. This is also very important for SMP since each core takes fetch bandwidth.
OTOH the lack of interlock was a kind of inverse: why waste transistors on something known at compile time anyway (same motivation for delay slots in branches, another idea that turned out not to be worth it). In reality, computation is dynamic, so runtime branch prediction (and later speculative execution) turned out to be a much better performance win, regardless of the design cost.
Edited to remove a comment about the 801, to which the author replied below.
I was a bit puzzled by the IBM 801 comment as the second para of the previous post mentions the 801 and links to an earlier post that is all about the 801. Was there something you think I missed in these earlier posts?
That didn't show as a link when I read it (shows as a link on my current device). Apologies for the oversight; I edited my comment.
The 801 was an insightful jump from then-current trends in processor design and that insight, to me, is what RISC is all about. There were lean and orthogonal instruction sets already (most notably, to me, the PDP-6/10 and Seymour Cray's work at CDC) but the central idea of offloading a lot of heavy lifting to the compiler (and recognizing Multics' insight of writing an OS in a HLL, pointing to a practically assembly-language free future) was groundbreaking, a kind of Special Relativity of computing.
Completely agree with your succinct summary of the what RISC is all about and the 801's place in the story. I've heard others call Cray's CDC's computers the first RISC designs, which I don't think is quite right. Probably deserves a fuller explanation than I can manage here though!
https://www.jwhitham.org/2016/02/risc-instruction-sets-i-hav...
> Nobody writes assembly code any more, except when they do,
His very first sentence turns out to be the primary assumption of RISC: essentially nobody writes assembly code so don't worry about making that easy, and depend on the HLL compiler to do a lot of heavy lifting. The small amount of assembly is just to boot the processor, boot a process, and some small glue in the OS, and as that stuff is heavily used but almost never written you don't have to worry about alleviating those peoples' suffering.
Oh, and compiler and debugger writers, and he does point out how some of these decisions make things more complex for them!
> CISC had always involved decoding instructions into micro-instructions.
That isn't actually true; into the 60s instructions were implemented in hardware, and the roots of CISC lie there -- essentially some quintessentially CISC instructions (BCD support, string handling, etc) were subroutines implemented in hardware because they were so common and that made the computer easier to sell. (as an aside: back in those days, when the machines had high level languages they were often unique to the vendor or even the specific machine model!)
Then again, even tiny machines like the 8080 that weren't really RISC or CISC had microcode, but that was later.
We are entering another interesting period with people building RISCV processors and SoCs on budget FPGA hardware. The distance between a working FPGA design and a working bit of custom silicon is a lot shorter than the distance between a simulated design and silicon. What I find amusing however is that we don't have a big "killer app" like the PC was at the time. We have cutthroat cost reduced embedded designs which are unforgiving.
Still, if you are someone who dreamed of designing your own bespoke processor architecture, now is a great time to be alive.
I think we do, though at a layer down the stack: do more at the edge using less power in the process. Just as the CPU<->memory pipeline lead to all sorts of interesting on-chip development (huge multi level cache architectures) network bandwidth in portable devices will increasingly be a bottleneck (it takes power to run the radio, and more and more devices will be competing for that bandwidth).
overall I think this means a lot less interesting work lower down in the stack
>Patterson and colleagues first looked to build a modified RISC design to run the Smalltalk programming language (in a project known as SOAR, for Smalltalk On A RISC) and then to form the basis of a desktop workstation (known as SPUR, for Symbolic Processing Using RISC).
"Symbolic Processing" here actually meant running Lisp: https://www2.eecs.berkeley.edu/Pubs/TechRpts/1986/6083.html"The restricted processor count also allows us to build powerful RISC processors, which include support for Lisp and IEEE floating-point, at reasonable cost."
"SPUR features include a large virtually-tagged cache, address translation without a translation buffer, LISP support with datatype tags but without microcode..."