## Why do Windows functions all begin with a pointless MOV EDI, EDI instruction?
[1]: https://devblogs.microsoft.com/oldnewthing/20110921-00/?p=95...
The fact that it’s xor rax, rax rather than xor eax, eax is also interesting as it’s one byte longer for exactly the same effect (modifying the bottom 32 bits of a register clears the upper 32 bits). It makes me think there’s something weird going on other than compiler stupidity. I’d be interested in seeing the code it was compiled from.
Though, I guess even if it was, it'd be silly to rely on it even on x86 only. Maybe it would still make for a nice fast-path? Dunno.
This just worsens my fear of changing "unnecessary" code when I don't know the original motivation for it.
AArch64 has a similar space: https://developer.arm.com/documentation/ddi0596/2020-12/Base...
And yes, PowerPC has a similar space as well holding hints like 'give priority to the other hardware threads on this core' and the like. https://utcc.utoronto.ca/~cks/space/blog/tech/PowerPCInstruc...
So consider the case of a standard mutex in the contended case. Normally the code will spin for a little bit before informing the kernel scheduler on the off chance that the thread that owns the lock is currently scheduled on another hardware thread. In that case it's in the best interest of the thread trying to grab the lock to shift most of the intracore priority to any other hardware threads so that it can potentially help the other hardware thread holding the lock get to a point where it gives up the lock quicker.
“random_samp_ele_crit=name
Specifies the random criteria for selecting the instructions for sampling. Valid values for this option are as follows:
ALL_INSTR
All instructions are eligible. This value is the default setting.
LOAD_STORE
The operation is routed to the Load Store Unit (LSU); for example, load, store.
PROB_NOP
Sample only special no-operation instructions, which are called Probe NOP events.
[…]”