Except in languages with a JIT compiler
Except in languages with a JIT compiler
Granted, the number of cases this distinction matters is relatively small, making a function faster only makes a program appreciably faster if that function is a bottleneck.
Erm, not sure how direct you want. But at least in dotnet you can use the Vector64, Vector128, Vector256, and Vector512 types [1] where each method gets effectively directly converted to a raw SIMD instruction. So, Vector512.LoadAligned call will be replaced with a raw SIMD register load instruction - supported by the CPU it is running on; and generally a JIT compiler will spot common patterns-of-use to optimise those too. There's no runtime check to see what is supported and no per-function-branching. It's as close to the CPU as you can get really (in a compiled language). Maybe I'm missing something?
If you just mean the difference between hand-coded assembly and the output of an optimising compiler, then sure, you can always be better with hand-coded assembly.
[1] https://learn.microsoft.com/en-us/dotnet/api/system.runtime....
[1] https://learn.microsoft.com/en-us/dotnet/api/system.runtime....
Azul and OpenJ9 additionally have server JITs, which widen the abilities of optimisations are available.
Additionally the ART cousin also does its own thing.
Each JVM implementation has its own approach how to do auto-vectorisation or mark intrinsic methods.
It is no different than talking about Ada, Fortran, COBOL, C, C++ and co compilers versus what ISO defines in the language standard.