Function multi-versioning in GCC 6
lwn.net
lwn.net
That said, I'm sure HPC folks compile with all the bells and whistles for the exact model, options, and probably CPU stepping enabled.
https://www.nas.nasa.gov/hecc/support/kb/preparing-to-run-on...
https://www.nas.nasa.gov/hecc/support/kb/broadwell-processor...
End user applications, that's another matter, and here using all the latest vector instructions etc. can make a difference. Usually less so than what one might hope, though. The really big deal tends to be using optimized libraries such as OpenBLAS, FFTW, MKL instead of doing numerical linear algebra yourself in a naive fashion, or using the reference netlib BLAS.
Another very common problem we see is poor application I/O patterns. Yes, every HPC site loves to brag how many GB/s their Lustre system does, but if you divide that by the number of CPU cores in a cluster, that ratio is quite low. Additionally, like other clustered file systems, Lustre metadata performance is relatively poor, so applications banging on lots of small files can easily tank the performance of the entire Lustre system.
(ICC solves this in a differently annoying way - where all the intrinsics are available, even if they are incompatible with the platform you are on.)
For some light entertainment, think about what happens with static initializers and compiling for different microarch flavours. If your C++ static init function happens to generate an AVX insn, and you've only got SSE2, welcome to SIGILL before main().
See for example stage 1 here: https://gcc.gnu.org/wiki/FunctionSpecificOpt (that document appears dated, but do things still work that way?) Afaik, clang/llvm have similar functionality.
Plus there are some intrinsics that are just macros, (sets, masks, etc), and you don't get them from the preprocessor just by setting the function target.
As an aside, that page really is dated - it is just early proposals afterall - as SSE5 didn't see light of day like that. VPCMOV ended up in AMD's XOP set.
ICC also has its auto as well as manual dispatch options:
auto: https://software.intel.com/en-us/node/682440 manual: https://software.intel.com/en-us/node/684505
I believe this is the area where Intel had their knuckles rapped for only working on "GenuineIntel" processors, and why there are big disclaimers on everything now. I've not tried using these myself as they aren't portable solutions.
Then again, perhaps that's all that's left for Intel to do now. Evolution of marketing ploys: transistors -> clock speed -> #cores -> instruction extensions.
Except for the fact that many of these new instructions perform some huge nontrivial operation in hardware that would've required hundreds or more regular instructions previously --- AES is a good example. It seems like a general principle that instruction sets tend to become more CISC-y over time, as dedicated hardware and instructions designed to operate on such replace slower software implementations.
Here is a pdf reviewing work on this[0], I guess. It is two years old though:
[0] http://llvm.org/devmtg/2014-10/Slides/Christopher-Function%2...
http://clang.llvm.org/docs/AttributeReference.html#ifunc-gnu...