CachyOS is a whole distro compiled with these flags, if possible, which is appealing.
CachyOS is a whole distro compiled with these flags, if possible, which is appealing.
[0] https://gcc.gnu.org/onlinedocs/gcc/Function-Multiversioning....
So it does have some limitations like not being inlined, same as any other external function.
A relocatable call within the same DSO can be a PC-relative relocation, which is not a relocation at all when you load the DSO and ends up as a plain PC-relative branch or call.
Ideally you should just multiversion the topmost exported symbol, everything below that should either directly inlined, or, as the architecture variant is known statically by the compiler, variants and a direct call generated. I know at least GCC can do this variant generation for things like constant propagation over static function boundaries, so /assume/ it can do the same for other optimization variants like this, but admittedly haven't checked.
You have bigger binaries, but the logistics are simplified compared to shipping multiple binaries and you should get the same speed as multiple binaries with fully inlined code.
Since they don't seem to be doing that, my question is: what's the caveat I'm missing? (Or are the bigger binaries enough of a caveat by themselves?)
It can be useful to duplicate the entire code for 8-bit vs 10-bit pixels because that does affect nearly everything.
error: inlining failed in call to 'always_inline' 'float _mm512_reduce_add_ps(__m512)': target specific option mismatch