Function multi-versioning in GCC 6 (2016)
lwn.net
lwn.net
If there is to be a CPU-specific dispatch mechanism, it should be more like a small binary/shell script that exec-s an entirely CPU-specific binary, dependent on runtime detection (cpuid, etc.). Dependent on your distribution mechanism, you might only download the appropriate binary. But if you do this, please remember to keep your binary sizes reasonable.
With C, C++, Rust and other such languages of course it's very difficult.
ETA: Also this excellent paper that is all about optimizing mem* notes that PLT cost adds up. https://dl.acm.org/doi/pdf/10.1145/3459898.3463904
For some clients it makes sense to statically link an implementation of these functions into your program. If you’re Google and your machines run one binary all day where you know the exact hardware they’re using, go wild with the assumptions that get baked into it. For everyone else, though, it is generally not the case that you want to reimplement libc, nor will you distribute separate binaries for each microarchitecure. In this case, you’ll pay the cost of linking against the system libraries anyways, and ifuncs give you a way to fold the cost of selecting hardware-specific optimizations into this process. This is why both glibc and libplatform do it. As far as I’m aware Microsoft chooses to actually put the check at the point where the optimized instructions would get used, but maybe that works better with how they do linking on that platform; I don’t know enough about it to have any insight to share.