Why did post-8008 CPUs not keep the on-chip stack idea?
retrocomputing.stackexchange.com
retrocomputing.stackexchange.com
Basically they have something like 256 registers set up so the output registers overlap the input registers of the next callframe. e.g.
iiirrrooo
iiirrrooo
iiirrrooo
Where i == input register, r == general purpose register, and o == output register. When code calls another function it places the first N arguments in the output registers. The called function reads arguments out of the input registers.This results in very high call and return performance. Until you exhaust the available register windows. At which point you fault, and the kernel has to manually copy the register window stack to memory. Similarly on return you may have reached the top of the stack you fault and the kernel has to copy the windows back from memory into the register windows.
For SPARC this was apparently an ok tradeoff as the "big iron" machines of the past generally did not recurse heavily.
Context switches also get more expensive, because you have to swap out the entire windowed register file, and not just what is visible to software. (That can be mitigated by additional hardware support, but the complexity is high.)
There not a huge amount of context switching (basically few processes relative to the number of cpus in the system), and known types of programs.
They knew in general most code they ran would not blow the stack. The cost of course if you ran atypical (for their designed purpose) your perf would be clobbered.
Likewise the ia64 registers were rotating, with a hardware stack engine.
To answer the main question: this wasn't abandoned in the industry. CPUs still do it. It just doesn't really provide much of a competetive benefit to architectures that do it, so it's largely been abandoned.
Where is the xtensa documentation? My random googling seems to indicate it has a more traditional fixed set of registers?
On SPARC for instance (super super super pseudo code here, I haven't written sparc assembly for... almost 20 years?)
mov %o0, an_argument
call a_function
...
a_function:
; the argument is now in %i0
or something like that.So it would be something like:
ld [%l0],%o0 ; Load whatever is pointed to by %l0 into %o0
call some_function ; Stores the old pc in %o7
; Result is now in %o0
some_function:
save ; Move the register window
; The argument is now in %i0
ld %i0,%l0 ; This does not overwrite the caller's l0
ret ; Return, essentially a jump to %i7
restore ; Called due to delay slot, restores the windowXtensa is simpler than Itanium, in that the windows are fixed 4-register sets and the spilling and filling is done in an exception handler. But it's more complicated than SPARC for sure. The details are in the "Xtensa Instruction Set Architecture Reference Manual", which is a PDF that is commonly available on the internet but not AFAIK actually distributed by Cadence itself.
"The reason" why memory is slower than registers (apart from size obviously) is that register dependencies are static and can be easily tracked by the core. Memory is way harder since you don't know all the addresses in advance (any other random store could actually be storing to the stack).
That's why they have to use costly (associative) structures like the store queue.
The way cores track the stack is usually for return address prediction because it's a huge win for little cost and almost no well behaved program overwrites return addresses manually (low mispredict rate).
As far as I know, all x86 cores emit 2 uops for a push or a pop (load/store + sp adjust). I guess it can save you some frontend bandwidth but that's about it.
I could have sworn I'd read something like this - the 'top' words of the stack always being kept in a very low-level cache - but I was unable to find a source on it.
> "I've found no particular evidence that the stack pointer was made a full 16 because they felt any need for a stack to be that large. It's clear that at least some experienced microprocessor developers (the MOS 6502 team) felt that an 8 bit stack pointer (256 byte stack) was plenty. It's possible that the 8080 designers disagreed, or it's possible that they felt they couldn't force a particular area to be RAM, as the 6502 designers could. (Even more than the MC6800, the 6502 design strongly encouraged page $00 to be RAM, so forcing page $01 to be RAM was no hardship.) Or perhaps it just didn't occur to them that registers pointing into memory could be any less than 16 bits."
To add a little bit of detail to the preceding paragraph:
A page in 6502 parlance is a contiguous block of 256 bytes in address space. Page $00 are the first 256 bytes of address space, page $01 are bytes $0000 to $00FF and so on. Page $00 (so called zeropage) is treated specially by the processor and is therefore required to be backed by RAM (not ROM or IO).
The 6502 has an 16 bit address bus but only an 8 bit stack pointer. The stack was fixed at 256 bytes (a page) in size and fixed in location at $0100-$01FF (page $01).
So the point made above is that because the 6502 already required the zeropage to be RAM it was easy to require the RAM for the stack in a fixed place too, while the designers of other contemporary processors didn't have that luxury.
Er, you mean page $01 is bytes $0100 to $01FF, right? And 6502 in general didn't use flat addressing, IIRC; eg `12FE,x` would address 12FE,12FF,1200,1201,... with increasing x register, rather than 1300,etc.
Yes, yes, thanks for the correction.
> And 6502 in general didn't use flat addressing, IIRC;
Not so sure about that.
> eg `12FE,x` would address 12FE,12FF,1200,1201,... with increasing x register, rather than 1300,etc.
This is correct but I think this is more due to the fact that the index register wraps around, and I would still call it flat addressing. A better argument to not call it flat would be the use of bank switching, which I think was common in 6502 based designs. But, yeah, this is probably just splitting hairs over terminology and I agree with you. Thanks for the corrections.
So I don't think it's an idea that's completely gone away, it's probably just that until relatively recently the complexity involved in doing it with modern code (with deep stacks with large stack frames) was not really a worthwhile use of die space. And now x86 and arm are so thoroughly dominant that other paradigms have difficulty getting traction.
[1] https://devblogs.microsoft.com/oldnewthing/20050421-28/?p=35...
Some of that is kernel prerogative. Still could be coded. Just avoid user code manipulating stacks. The introduction of instructions to support those features is conceivable. Still while avoiding "unrestricted addressing of the entire call stack at will".
That's kind of the whole point - call/return can be vastly more secure than 'part of a data space' as it is now. Benefits would accrue. And challenges.
Some features may be inherently dangerous and not supportable on a secure processor.
The hardware stack enabled nice tricks back in the day, at the whole CPU was a fun little beast.
4-bits bus and addressable units (nibbles), plenty of 64bits registers, 20 bits addresses...
The problem is that it does not really fit to complied code. The compiler usually has no clue how deep the stack is when a function is called (esp. with function pointer, virtual methods, etc), so it will not or can not generate optimal code for this. Plus you get the trouble with interrupts.
However, modern CPUs understand how the stack is used and actually will maintain and on-chip stack for you. That is, if you are using the standard function prologue and you dont play any tricks with the stack-modifying instructions.
Because on later CPUs can exhaust it with few calls.