Unfortunately you can't get away with just one set of spare registers on x86 interrupt handling. Interrupt can be interrupted by another interrupt. So practically they have to save registers anyways. Unless there'd be one set for each level...
So those "10 extra registers" would be practically useless.
Having written kernel device driver interrupt handlers, I can also say CPU time is not spent saving and restoring registers, but waiting for glacially slow PCI-e MMIO register loads and stores (don't have exact figures at hand, but I remember one access can take 300-800 nanoseconds, that's thousands of CPU clock cycles). During one such fetch you might be able to store and restore registers tens, if not hundreds of times. X86 interrupt handling is a bit like stopping a freight train to pick up a single letter.
However, that kind of feature can be very useful on a microcontroller to help lowering interrupt handling latency.
Only when an interrupt actually does get interrupted. You could make the common case faster. Not that it matters if PCI-e is the bottleneck.
So what do you suggest could be done?
There's the cost of a branch, but you could remove that by making it a hardware feature. It's a branch that should predict well though.
Reading and writing to memory mapped I/O registers is how you send commands to peripherals, read their status, etc. These registers are not really memory at all, but representation on various logic states on the peripheral. So what you write is often not what you read back.
Note that I/O memory mapping has nothing to do with usermode memory mapping.
When an interrupt or system call happens, all the registers get pushed to the stack. The stack is typically in L1 cache (and subject to various CPU optimizations), so it's really fast to push the 16*64 bit registers to the stack.
System calls, interrupts and context switches are "slow" not because they have trivial overhead like pushing registers. What's really consuming the time is secondary effects like TLB flushes, changes to the page tables, polluting the branch predictor, etc. It takes a very long time for the CPU to "warm up" again after a context switch.
I've made the remark because it's fairly common to mistake mode and context switches.
it's really fast to push the 16*64 bit registers to the stack
Since the CPU has 180 registers (with only 16 names), why don't we need to push all 180 to store context?No, it is not.
Modern CPUs all use Physical Register Files or PRFs for implementing OoO. The way they work is that there is one backing register file of ~150-200 registers, and a naming table in the frontend. Each register file in the backing table can be in one of 3 states -- waiting for data, contains data, or clean. Every time an instruction that writes data to a register is executed, during the register rename stage the renamer picks one register from the clean set, assigns that as the output of the uop and the current value of the logical register, and sets it as waiting for data.
That is, let's say you execute:
add RAX, RAX
and the current value of RAX in the rename table was PRF#1, with PRF#1 carrying data and all other physical registers clean. In the renamer, this would then turn into:
add PRF#2, PRF#1, PRF#1
if this instruction was immediately followed by another add, RAX, RAX, it would then turn into:
add, PRF#3, PRF#2, PRF2, and after renaming that instruction RAX points to PRF#3.
and so on. In PRF machines, architectural registers only exist as pointers into the PRF.
The reason for this design is that it makes OoO execution simple, and reduces unnecessary data movement. An instruction is ready to execute once all it's inputs have data in the PRF, and when it executes there is no need to move data into some special architectural register.
Registers are reaped from the "haves data" set into the clean set once no instruction or architectural register points at them. Note that multiple architectural registers can point to the same physical register: mov rax, rbx is resolved in the frontend just by pointing rax to the same register as rbx, and there is typically a special register for the value 0 that all common register-clearing operations use.
The really short explanation about the forwarding network is that writing/reading the PRF takes too much time for you to read it and do work on the same cycle. If you operated on it alone, all operations that were dependent on each other would have multi-cycle latencies.
Since this sucks, in addition to the PRF the machines have a forwarding network, where the result of every instruction is broadcast to all the execution units of a similar type. If the broadcasted PRF number matches the register you want, it's read from the forwarding network instead of the PRF.
AFAICT, the Mill works by completely eschewing the final register file, and exclusively using the forwarding network with some local storage at each execution unit for the past values on the network.
PRF is Physical register file? As opposed to the logical which can point to different backing store(the actual SRAM.)
Register renaming and OoO operations are really enabled by micro-ops? In other words on a CPU with completely hard-wired control unit such things would never be possible.
Yep, it's the actual renamed registers.
> Register renaming and OoO operations are really enabled by micro-ops? In other words on a CPU with completely hard-wired control unit such things would never be possible.
No, they're orthogonal concepts. The first OoO processor (the System 360 model 91) didn't have uops.
Yep. PRF means the single backing store where all values go, and PRF systems are named for it because they are usually contrasted to the other very common way to implement OoO, ROB-backed systems, where there is a specific location for the final architectural value, and possibly multiple in-progress values in the ROB. Intel used values in the ROB in all their CPUs up to Sandy Bridge, which was their first PRF design.
(The main difference between the types is that ROB-backed is faster in circuit delay, but moves more data than a PRF design, so on the same process they potentially clock higher but produce more heat. PRF-based designs are also easier to scale wider than the ROB-based design.)
> Register renaming and OoO operations are really enabled by micro-ops? In other words on a CPU with completely hard-wired control unit such things would never be possible.
No, you can do renaming and OoO well without uops, it just requires that your instruction set is designed to be amenable to it. Before x86 had uops, it couldn't do OoO and the competing RISC CPUs were way faster because of this. PPro added uops precisely because they made it possible for an x86 cpu to do OoO, and eventually x86 beat all the competing RISC chips at their own game.
(Exposing more virtual registers have two main costs, then: you lose backwards compatibility for all programs that use the extended register set, and to back them up with the same technique you need several more real registers.)