Hennessy and Patterson win Turing Award
acm.org
acm.org
Before that, instruction sets were driven more by aesthetics and marketing than performance. That sold chips in a world where people wrote assembly code -- instructions were added like language features. Thus instructions like REPNZ SCAS (ie, strlen) which was sweet if you were writing string handling code in assembler.
H&P must have been in the queue for the Turing award since the mid-90s. There seems to be a long backlog.
REP MOVS is still the fastest and most efficient way to do mempcy() on most Intel hardware, and likewise REP STOS for memset(), because they work on entire cachelines at once.
It's worth noting that, if it weren't for a brief period of time during the late 70s/early 80s when memory was faster than the processor, RISC as we know it may have never been developed; the CISCs at the time were spending the majority of the time decoding and executing instructions, leaving memory idle, and that's what the RISC concept could make use of --- by trading off fetch bandwidth for faster instruction decoding and execution, they could gain more performance. However, the situation was very different after that, with continually increasing memory latencies and now multiple cores all needing to be fed with an instruction stream putting RISC's increased fetch bandwidth at a disadvantage. Now, with a cache miss taking tens to hundreds of cycles or more, it seems a few extra clock cycles decoding more complex instructions to avoid that is the better option, and dedicated hardware for things like AES and SHA256 obviously can't be beat by a "pure RISC". It was almost "game over" for "RISC=performance" in the early 90s when Intel figured out how to decode CISC instructions quickly and in parallel with the P5/P6, and rapidly overtook the MIPS, SPARCS, and ALPHAs that needed more cache, higher clock speeds, and power consumption to achieve comparable performance.
Certainly, it makes one wonder whether, had that brief moment in time not existed and memory was always significantly slower than the processor, would CPU designs have taken a completely different direction?
Would that be the original REP MOVS/STOS, or the "fast strings" (P6)? Or the "this time we really mean it fast strings" (Ivy Bridge)? Or the "honestly, just trust us this time fast strings" (Ice Lake)?
> Now, with a cache miss taking tens to hundreds of cycles or more, it seems a few extra clock cycles decoding more complex instructions to avoid that is the better option
You can have both, actually. E.g. RISC-V with the compressed instruction extension achieves higher code density than x86-64.
> and dedicated hardware for things like AES and SHA256 obviously can't be beat by a "pure RISC"
Well, if you're really looking for minimal instruction sets, RISC is way bloated; IIRC single-instruction computers can be Turing complete. Obviously they are not very useful in practice. I think a better approximation of the RISC philosophy is "death to microcode", that is, the instruction set should match the hardware that the chip has. So if your chip has dedicated HW for some crypto or hashing algorithm, I wouldn't consider it "un-RISCy" to expose that in the ISA.
I believe most libc implementations are similar, such as glibc: https://sourceware.org/git/?p=glibc.git;a=blob_plain;f=strin...
No real computer is Turing-complete, since it has only a finite amount of memory.
Having said that, Hennessy and Patterson definitely deserve the award.
The software equivalent is to forget about the distinct human-invented concepts of programs, input data, memory representation of data, and languages - it can all be optimized together with global visibility to all aspects. Like partial evaluation optimizations, without the partial part.
Our compiler technology still has ways to go with this one.
I enjoyed that it was a simpler read then a lot of the circuits-type of books that are part of a EE/CE curriculum, but I always felt there was this lack of "hard science"/physics in the book.
And perhaps it was just not a topic they felt fit with the vision of what this book is suppose to be, and it likely came to be a better decision to abstract that part away for readability.
It's a long time since I read it, but from my memory that's the kind of thing the book is about. I'm not sure how you'd determine those kinds of things in a more "hard science"/physics fashion.
I would recommend COaD to any beginner who wants to learn some basic concepts of computer design (and if working with FPGAs why not build one).
CA deals with more advanced concepts but doesn't overwhelm you with math and circuit theory (as you noted) so it's a natural progression (from COaD). I think something more advanced and "hard science" should be part of a post-graduate curriculum
But man, 6 months out and I retained nearly nothing. It's just so esoteric and not very relevant to my every day schoolwork(then) and career(now) that it's slowly seeped out of my brain.
Given that most of my career has been pretty far removed from the compiler/assembly instructions, much of the content wasn't directly relevant to what I've been doing, but I still find myself using many of the tools that are outlined in the first chapter ("Quantitative Principles of Computer Design") when dealing with performance problems or analysis in high-level-language-world, which happens pretty often.