It’s my stack frame, I don’t care about your stack frame
blogs.embarcadero.com
blogs.embarcadero.com
The caller ensures that the stack is 16-byte aligned at the point of the function call.
“At the point of the function call” means just before the CALL instruction is executed, the top of the stack. No further explanation is given, but an inference can be made: the Darwin kernel is using SIMD instructions to quickly move data on the stack, possibly during context switches.
http://pages.cs.wisc.edu/~weinrich/papers/method_dispatch.pd...
Regardless of the rationale, it really isn't a big deal at all. He'll need to jump through the stack alignment hoops anyway when he gets to finishing his 64-bit compiler, and so what if he has to do the same with the 32-bit compiler?
I agree that the ABI could use the convention that the callee sets ESP to have whatever alignment it wants, if needed for its SSE ops or whatever. But really, why care? Sure, it'd be interesting to hear the rationale... but in the end, we just have to follow the ABI, and mandating 16-byte stack alignment isn't all that crazy.
Consider a call to "void f(long x)": this means very different things for i386 code (long is 32-bit, little-endian) and Mac powerpc code (long is 64-bit, big-endian).
And you don't have to insert extra code at all call sites. You can assume that your stack is correctly aligned at entry too, so if you just make your stack frame a multiple of 16 bytes, you'll be fine.
You're also assuming that there is no intra-procedural stack modification that may be in effect while a function is being called.
The stack is not aligned at function entry because the CALL instruction pushes the address on the stack, thus having alignment at the call site forces the stack to be unaligned at function entry, however, it is unaligned by a known amount so you can produce aligned addresses by subtracting a constant value (skipping the BAND instruction). The featured rant misses this point, you get aligned addresses just as easily at the bottom of your frame as at the top.
Realigning the stack at runtime though, is expensive and complicated.
Or is there some trick that I'm missing?
Edit: You're right, optimized code might not save the frame pointer and this wouldn't provide any benefit if your frame had unknown size. But usually the preamble saves the old ESP in EBP and it is restored at the end. Maybe there is another alternative.
I guess the answer is that realigning at runtime costs a register, but most code has to pay that cost anyway.
There's no specific timeframes (and it looks a bit outdated), but here's the latest roadmap: