In point of fact beyond the instruction decode stage all modern cores look more or less identical.
In point of fact beyond the instruction decode stage all modern cores look more or less identical.
The divergence was one of philosophy, and had unexpected implications.
CISC was a “business as usual” evolution of the 1960s view (exception: Seymour Cray) that you should make it easy to write assembly code so have lots of addressing modes and subroutines (string ops, BCD, etc) in the instruction set.
RISC was realizing that software was good enough that compilers could do the heavy lifting and without all that junk hardware designers could spend their transistor budget more usefully.
That’s all well and good (I was convinced at the time, anyway) but the results have been amusing. For example some RISC experiments turn out to have painted their designs into dead ends (delay slots, visible register windows, etc) while the looseness of the CISC approach allowed more optimization to be done in the micromachine. I did not see that coming!
Agree on the point that the cores themselves have found a common local maximum.
In the 70s, everyone designing an ISA was doing CISC. Then in the 80s, everyone suddenly switched to designing RISC ISAs, more or less overnight. There weren't any holdouts, nobody ever designed a new CISC ISA again.
The only reason why it might seem like there was a divergence is because some CPU microarchitecture designers were allowed to design new ISAs to meet their needs, while others were stuck having to design new microarchitecture for legacy CISC ISAs which were too entrenched to replace.
> For example some RISC experiments turn out to have painted their designs into dead ends
Which is kind of obvious in hindsight. The RISC philosophy somewhat encouraged exposing pipeline implementation details to the ISA, which is a great idea if you can design a fresh new ISA for each new CPU microarchitecture.
But those RISC ISAs became entrenched, and CPU microarchitecture found themselves having to design for what are now legacy RISC ISAs and work around implementation details that don't make sense anymore.
Really the divergence was fresh ISAs vs legacy ISAs.
> while the looseness of the CISC approach allowed more optimisation to be done in the micromachine.
I don't think this is actually an inherent advantage of CISC. It's simply result of the shear amount of R&D that AMD, Intel, and others poured into the problem of making fast microarchitectures for x86 CPUs.
If you threw the same amount of resources at any other legacy RISC ISA, you would probably get the same result.
I don't see how that follows at all? MIPS in fact is about as pure a "RISC" implementation as is possible to conceive[1], and it shares all its core ideas with RISC-V. You absolutely could make a deeply pipelined superscalar multi-core MIPS chip. SPARC has the hardware stack engine to worry about, but then modern CPUs have all moved to behind-the-scenes stack virtualization anyway.
No, CPUs are CPUs. Instruction set architectures are a vanishingly tiny subset of the design of these things. They just don't matter. They only seem to matter to programmers like us, because it's the only part we see.
[1] Branch delay slots and integer multiply instruction notwithstanding I guess.
And, yes, they learned the hard way that register windows are a bad idea. Patterson says they did it because their compiler didn't have as good register allocation as Stanford's compiler.
Am I correct in recalling they removed branch delay slots in a later iteration of the chips?
Had to write MIPS assembly by hand recently, incredibly counter intuitive.
An internal fixed with encoding inside the instruction cache may work.
Uh, why? Instructions start at byte boundaries. Say you want to decode an whole 64 byte cache line at once (probably 10-14 instructions on average). Thats... 64 parallel decode units. Seems extremely doable to me, given that "instruction decode" isn't even visible as a block on the enormous die shots anyway.
Obviously there's some cost there in terms of pipeline depth, but then again Sandy Bridge showed how to do caching at the uOp level to avoid that. Again, totally doable and a long-solved problem. The real reason that Intel doesn't do 10-wide decode is that it doesn't have 10 execution units to fill (itself because typical compiler-generated code can't exploit that kind of parallelism anyway).
Overlaid on a higher resolution die shot https://static.wixstatic.com/media/5017e5_982e0e47d7c04dd693...
Here's a POWER8 floorplan: https://cdn.wccftech.com/wp-content/uploads/2013/08/IBM-Powe... (The decoder is subsumed under IFU)
And POWER9 https://www.nextplatform.com/wp-content/uploads/2016/08/ibm-...
Didn't find any recent, reliable sources for Intel cores.
I wonder how wide SIMD has to get before you treat it like a CPU embedded into cache memory.
Though I guess we are already looking at SIMD instructions wider than a cache line…