The mysterious second parameter to the x86 ENTER instruction
devblogs.microsoft.com
devblogs.microsoft.com
I suspect it is completely microcoded and there is no dedicated silicon to implement it (well other than ROM space to store the microcode I guess).
From everything I've seen, the x86 and x86-64 instruction sets are loaded with "it seemed like a good idea at the time", "if you couldn't count on FPU / MMX / SSE instructions being available...", and similar waste-of-silicon instructions.
Plus, there's a feedback loop: Compilers writers restrict themselves to subset S of the full instruction set for cost-benefit reasons <==> CPU designers devote minimal time / effort / silicon to improving the performance of non-S instructions.
They are, but a lot of "instructions" these days don't even map onto native silicon logic anyway like they used to. It's all microcode these days.
I know we're splitting hairs here, but really the only difference between microcoded instructions and "native" instructions is that the instruction decoder directly produces u-ops versus doing a micro code table lookup. The microcode table lookup just ends up producing u-ops, and those all go through the rest of the out-of-order engine just the same. It's really just the frontend that is the difference. For CPUs that have a trace (u-op) cache, I think the results of microcoded instructions can also be stored there.
Of course the ability to modify non-microcoded instructions is significantly smaller and limited to what was planned for at design time.
A silicon vendor I worked at called them CYA (Cover Your Ass) bits :)
This is all the very small cost of backwards compatability.
https://github.com/ianopolous/JPC/blob/master/src/org/jpc/em...
https://github.com/ianopolous/JPC/blob/master/src/org/jpc/em...
"JPC is a fast modern x86 PC emulator capable of booting Windows up to Windows 95 (and windows 98 in safe mode) and some graphical linuxes. It has a full featured graphical debugger with time travel mode along with standard features like break and watch points."
1> (let ((*opt-level* 0))
(compile-toplevel '(let (a) (let (b) (let (c) (list a b c))))))
#<sys:vm-desc: 9fb6870>
2> (disassemble *1)
data:
syms:
0: list
code:
0: 04020001 frame 2 1
1: 04030001 frame 3 1
2: 04040001 frame 4 1
3: 20030002 gcall t2 0 v00000 v01000 v02000
4: 08000000
5: 10000C00
6: 10000002 end t2
7: 10000002 end t2
8: 10000002 end t2
9: 10000002 end t2
instruction count:
8
#<sys:vm-desc: 9fb6870>
The frame 4 1 instruction means we are at depth 4, and there is 1 variable here.One of my favourite examples is the x86 task switching support. The 386 had everything needed to implement context-switching and multitasking -- in hardware. (That notion was in vogue, once. See the ICL 2900.) If you set up the right structure in memory, and turn on the hardware task switching, x86 processors will automatically context switch all of the registers in and out of that in-memory structure, at the appropriate privilege boundaries (interrupt, return from interrupt, etc.) So your kernel can switch between processes like this (more or less):
mov next_proc_ptr, tr
iret
It also provides IO port protection - there's another table that lets you specify which ports the process can access, which the processor will examine, when an IO instruction is executed in user mode, presumably for something microkernelish.Now that's the good ol' CISC spirit. It's quite a complex feature and yet I'm not sure any OS ever used it.
It's certainly much faster today to use sequences of mov instructions. In practice, it's a block move of multiple registers to memory/cache. That's something any modern pipelined and microcoded processor will have specialized support for. What instruction or instructions generates that internal action, is largely beside the point.
This is no longer so true today, however the Linux kernel for example doesn't use any floating point or vector instructions, in order to avoid having to save these registers on every system call or interrupt.
The real inflexibility was scheduling (see my other comment)
it doesn't use them by default, but code that needs the FPU can call kernel_fpu_begin() to save those registers if needed. https://yarchive.net/comp/linux/kernel_fp.html One big user is the AMD graphics driver: https://lore.kernel.org/amd-gfx/ee26b06b-f1a0-5a8f-388c-13c1...
One might think that since interrupts can switch tasks, you could set up a simple round-robin scheduler by hooking this mechanism up to the timer interrupt somehow, or have device drivers in separate tasks that become active when receiving an interrupt.
But a "task state segment" on x86 is actually more like a pre-allocated stack frame, with interrupts creating a linked list of these. A task can't be reentered while it is active, and it is supposed to exit by returning to the previous task that was interrupted. Anything else requires software scheduling.
The only reason to ever use this mechanism is for handling exceptions on a specially reserved stack (for example when the kernel stack overflows).
And then you get the inflexibility: There's a limited number of hardware task slots; I think you can pick which registers to save (or rather how many to save), but you always save all of them: a kernel syscall might be more judicious and only save the registers it uses (of course saving all registers if returning to a different task); it's hard to change because software must run on old and new cpus; etc.
There's exceptions to the rule of thumb. Rep mov can be faster than doing it yourself. It's also quite possible for a combined operation to do two things at the same time rather than one after the other.
When nobody uses such a feature is typically either because it has worse performance then alternatives or some missing feature. There have been many such cases in x86.
They did.
Linux used x86 hardware task switching and TR (task register) for years. See "struct tss_struct" and "switch_to" in older kernels.
Linux also allowed userspace processes to get hardware-managed access to individual I/O ports using the per-task x86 I/O port bitmap, and to local segments using the x86 LDT which were useful for running 16-bit code in a 32-bit task.
The hardware task switching was eventually changed to software task switching. By that time, software switching was faster, and it was always more versatile. This change removed the hardware-imposed limit on the number of processes and threads Linux could support.
Here's a nice article about the evolution of task switching in Linux. https://www.maizure.org/projects/evolution_x86_context_switc...
TFA gives at least a partial answer to my earlier q: https://news.ycombinator.com/item?id=38544640
I think I misunderstood the blog author's question. He wasn't asking about the ENTER statement with a non-zero argument, but about the special ABI with arrays of frame pointers.
[1] https://en.wikipedia.org/wiki/Nested_function#Languages
[2] https://www.gnu.org/software/c-intro-and-ref/manual/html_nod...
There's a variety of implementations there though - some seem to be talking about closures, and at least one (Rust) doesn't appear to support parent variable access.
Isn't this just a lexical closure? It's common in scripting languages.
In Pascal local functions are not closures and do not affect variable storage. All variables are still stored in the stack frame (unless optimized to be in registers), and local functions can access their parent's (& grandparent's etc) variables via pointers to their stack frames, which are normally held globally in an array of frame pointers called the display.
For a Pascal depth-N nested local function to access it's parent's variables it would use the depth-(N-1) display entry, for it's grandparent's variables the depth-(N-2) entry etc. When a parent calls a nested function the only thing that needs to change in the display is to set the depth N-1 display pointer to it's own stack frame, and to replace it with the old N-1 pointer when the call is complete. The grandparent/etc entries don't need to change.
This x86 ENTER instruction appears to be designed for a more general case than Pascal, since it's apparently copying the entire display (array of frame pointers), not just doing the more efficient incremental update/restore that is all Pascal needs.
It's interesting to see Raymond Chen writing about this "almost always unused" ENTER parameter since it suggests that compilers are not using it even for those languages that might seem to be able to make use of it. Perhaps, as for Pascal, it's not the most efficient way, or perhaps just because languages such as JavaScript don't normally compile to machine code.