X86-64 is not under heavy register pressure strain. That idea is a legacy from x86 (plain, 32b).
x86-64 has the same register count as ARMv7 and few bothered disabling the frame pointer there, even though it’s a load-store architecture.
X86-64 is not under heavy register pressure strain. That idea is a legacy from x86 (plain, 32b).
x86-64 has the same register count as ARMv7 and few bothered disabling the frame pointer there, even though it’s a load-store architecture.
This makes the case where that one extra register name makes all the difference much rarer, arguably turning it from "I demand a compiler flag" to "Let's just hand-write the machine code for this one very special routine if our performance data suggests it's worth it".
† Internally a modern CPU has far more actual register, to enable a feature called "register renaming". But we can only talk about them using their canonical names, and x86-64 adds eight more of those.
But i agree the impact of preserving frame pointers is generally quite small and doesn’t often actually need mitigation - on amd64 there’s not much impact from losing 1 more of 16 arch registers.
https://www.agner.org/forum/viewtopic.php?t=41
Intel Alderlake has performance events for tracking it:
https://github.com/intel/perfmon/blob/974c69919b2a9dfd8278cf...
But even before this you had store to load forwarding on x86. I'm not saying you have, but before inventing a performance problem it is worth spending time trying to diagnose it with thorough profiling (e.g. [1]). The Fedora frame pointer patch did a thorough performance analysis and performance will be revisited again. Unfortunately there are a lot of arm chair performance experts who haven't spent time looking into the details.
[1] https://perf.wiki.kernel.org/index.php/Top-Down_Analysis