A Brief Retrospective on SPARC Register Windows
danielmangum.com
danielmangum.com
This extra hardware only helped that one problem, which is pretty wasteful.
Instead, more modern CPU architectures have focused on implementing features that improve performance of the whole memory subsystem.
Even a simple on-die data cache will go a long way to solving the memory bandwidth, as most spills and refills will hit the cache. But modern out-of-order designs have store buffers that will detect when the same address is stored, loaded and overwritten in quick succession, allowing the spilled value to avoid even hitting the L1 data cache. I believe some architectures (like AMD's Zen 3) can even eliminate the memory operation entirely, by detecting pairs of stores and loads and redirecting them to renaming registered.
The other issue caused by spills is the extra instructions themselves. Some architectures (like arm) implemented load/store multiple instructions that more or less solved this problem, but others just relied on extreme out-of-order execution to improve throughput of all instructions.
The article concluded that the idea of Register Windows went away because it was a bad idea that didn't really work in practice. But I think that's only half the reason, and that could have probably been fixed with some design iteration. The other half is that idea of Register Windows were simply unneeded after various improvements to the memory subsystem in general.
The most important of those features is register renaming. (You mentioned it briefly; the article doesn't touch on it at all.) When your CPU actually has a few hundred hardware registers -- which many modern CPUs do! -- the fact that your ISA only has names for a dozen or so of those suddenly becomes much less important.
> Some architectures (like arm) implemented load/store multiple instructions that more or less solved this problem
The LDM/STM instructions were removed in A64 (64-bit ARM), actually! They were clever at the time, but they're difficult to work around in the pipeline.
EDIT: given that superscalar implementations already implement register renaming logic
The Itanium also got register windows but designed differently. First, the window size was variable, so that a function didn't have to allocate registers it didn't use.
It also used a separate protected register stack. The architecture allowed registers to be spilled/filled automatically in the background in-between load and store-instructions instead of all at once, but I'm not sure that any implementation did that.
BTW. The return pointer was also on the register stack, so it couldn't be overwritten by a buffer overflow.
I remember having to delve into a bug in some C code and GDB was acting absolutely strange. So I had to drop into assembly level debugging and was completely confused as to why a register was suddenly changing.
That led to me pulling out a big book on SPARC and I had to learn about register windows for the first time which explained everything and helped me solve the bug.
One of my prouder problem solving experiences considering I knew very little C at the time, even less about assembly and nothing about SPARC.
It’s a shame AMD abandoned it in its failed quest for x86 dominance.
Also, would register windows have caused speculative execution problems analogous to "Spectre"? You did have to keep some weird number of bytes (96?) available for register spills when you created a new stack frame. It might be possible to get Spectre like leaks from the kernel in that spill zone.
Unlikely, they optimized for the wrong case, and made multitasking context switches very painful. From the kernel's perspective, you could add hardware workarounds, but they would be unwieldy and likely hard to code securely around.
> A modern x86_64 chip is immensely complicated internally.
Any OOO chip is going to be. x86 has somewhat more complicated instruction decoders due to variable length instructions, but past that, it's not very much more or less complicated than any other contemporary ISA.
> would register windows have caused speculative execution problems analogous to "Spectre"?
Any processor with a cache that is modified by speculative execution is vulnerable to Spectre.
So, the "lack of complexity" comes from a "lack of sophistication." There are no free lunches is the point and you're going to end up with a chip that's just as complex as any other OOO out there, and very little of that complexity is going to be due to the ISA, once you match performance.
Depending on how liberally you want to define 'register windows', particularly if you include "two register sets", one could certainly say this is true. Many architectures have dual register sets, usually touted as for "fast interrupt handling" or other optimization based on not having to save the whole register file. Even the venerable Z80 has something like this. I have always assumed that's where the original idea grew from: if being able to speed things up by not push/pop-ing the registers is good for one type of context change, why not all/more of them?
I'm not enough of a theoretician or pedant to augur where register windows begin or end, however.
Of course, you would be storing a lot of garbage, but its a linear and trivially parallel regarding other load and store tasks.
Quoting WiKipedia:
The original Sun-4 series were VMEbus-based systems similar to the earlier Sun-3 series, but employing microprocessors based on Sun's own SPARC V7 RISC architecture in place of the 68k family processors of previous Sun models.
These all proceeded the 4/60 aka SPARCstation (1), which wasn’t introduced until April 1989.