x86 is fundamentally a CISC; if you treat it like a RISC, it will definitely disappoint.
x86 is fundamentally a CISC; if you treat it like a RISC, it will definitely disappoint.
I asked around at the time and someone mentioned that I might have overtaxed certain execution ports or something like that, but yeah that just cemented my belief that x86 optimization is not my cup of tea anymore. Better to spend time learning how to write code the compiler can optimize well.
These things are knowable if you have enough curiosity and maybe masochism :)
Nowadays, counting cycles is a game for monks who don't actually have to get anything done.
(But CPUs are usually better than you think they are.)
That’s only part secrecy and part to give them freedom to change it. It is of course somewhat described in their patents.
I honestly don't know if it's worth it to try to optimize branch prediction in compilers these days, beyond the obvious step of putting the highest probability target next (for fallthrough prediction) and generally laying out hot parts of the code together. TurboFan and most other dynamically-optimizing compilers put rare code at the end of functions, and that's a huge boost.
I'm wondering how the incentives play out to keep this stuff private?
https://www.intel.com/content/www/us/en/developer/articles/n...
I think software is not a huge profit center for them.
The original comment presents:
> The actual details there are [1] too secret for Intel to want to accurately describe them in gcc, [2] they’re different across different CPUs, [3] and compilers just aren’t as good as you think they are.
2 and 3 could just be the whole story. Although we haven't actually accumulated any evidence here for 3, given that the original story was about surprisingly getting beat by a compiler, despite performing a seemingly obvious optimization.
Actually, compiler optimizations like scheduling tend to be neutral to negative on x86 because they increase register pressure. You’d probably want to do “anti-scheduling” and hope the CPU decoder takes care of it if anything.
Interesting that it generates suboptimal code for non-Intel processors.