Moreover, can you influence GCC over what instructions it emits? On Gentoo for example, you can setup your make file so that the while system is compiled with your specific processor in mind.
Moreover, can you influence GCC over what instructions it emits? On Gentoo for example, you can setup your make file so that the while system is compiled with your specific processor in mind.
For what it's worth, this is largely unnecessary nowadays. Tuning for a specific processor used to allow the compiler to use instruction scheduling and enabled use of SSE instructions. However:
1) Modern x86 CPUs all use out-of-order execution, making instruction scheduling unnecessary. (Some older Atom CPUs are in-order, but their performance is pretty awful even with instruction scheduling.)
2) If you're building for x86_64, you can assume a certain baseline of instruction set support -- in particular, CMOV, SSE, and SSE2 instructions are always available on any 64-bit CPU. While these don't cover all instructions, it does cover the bulk of what the compiler is likely to want to generate in typical application code.
Not true, compilers remain extremely conservative in their default codegen, we’re talking pre-nehalem (SSE2 at the top end).
I just checked on godbolt, by default u32::count_ones compiled with rustc 1.62 with -O generates 15 lines of assembly hand-counting the bits in the word.
Adding `-C target-cpu=x86-64-v2` reduces that to a single popcnt.
Even if you don’t care for SSE3+, there’s a handful of really useful instructions which are not part of the baseline x86 profile, especially for bit-twiddling and endianness manipulation: popcnt, lzcnt (v3), movbe (v3), BMI1 and BMI2 (v3)
Otherwise I agree that these are awesome, I know a fast Hilbert-curve implementation using those.
I'm not aware of an other extensions with such a slow implementation in first generations.
[1] https://www.agner.org/optimize/instruction_tables.pdf
It's unfortunate that PDEP and PEXT is part of BMI2. AMD probably wanted to push BMI2 for the other instructions and they had to include emulated PDEP and PEXT before they had a proper implementation for those.
In short, the out of order instruction buffer can do some amazing stuff for code that would otherwise run much slower, that doesn't mean you can't gain or lose performance by reordering instructions. For non-trivial code the best composition is almost certainly different between CPUs.
Lots of high performance microarchitectures use assumptions about control flow for optimisation e.g. use the direction (backward, forward) of conditional jumps and branches as weak prediction, convert conditional branches over a single instruction into predicated execution, fuse common instruction pairs into a single micro-operation. At least limited cooperation between compiler and CPU core is required to make good use of those features.
It would definitely have to include the various variations. I ran it on my system, and of the top 38 instructions, 13 of them are variations on `mov`. The 22nd most common instruction is `and`, 26th is `shl`, 34th is `shr`, 36th is `or`, 39th is `imul`. You need to know those.
/usr/local/bin with tuned CFLAGS sounds better for such a statistical analysis.
Yes, of course. On a Gentoo ~amd64 system where everything is compiled with "-O3 -march=native", I get 1311 different ASM instructions: https://gist.github.com/stefantalpalaru/932b6d0ef439756e825d...