However, most SSE2 instructions won't be emitted by GCC until you crank things up to -O3.
However, most SSE2 instructions won't be emitted by GCC until you crank things up to -O3.
In fact, x86_64 specifically guarantees the availability of SSE2.
Vectorising non-vector C code is difficult (and likely relies on other expensive optimisations around flow control manipulation).
And it’s not a “sure fire” optimisation because the compiler can absolutely vectorise a loop whose 99p is an iteration or two making the vectorisation a complete waste of compilation and runtime.
So being an expensive and not perfectly reliable optimisation, it makes sense for it to be at O3.
That has nothing to do with ISA subsetting/compatibility.
These will be used by default (at O2 without vectorization enabled) for single-precision floating point math on x86-64 (where stack-based FPU is deprecated).
Vector instructions you're referring to end with "ps", they're very unlikely to be emitted for typical fp operations.
It’s not universally good. It’s just that if you take naïve C code that does bulk operations on arrays, and compile it at -O3, you might get a very nice improvement from the auto-vectorizer. With some additional restrict keyword annotations, you can sometimes improve the performance significantly.
Then there are non-standard tricks you can use, like __builtin_assume_aligned().
This is something of an arcane art. Sometimes it just doesn’t work at all, sometimes it’s not obvious how to write your code so that the autovectorizer can work well, and getting the autovectorizer to work requires enabling various code transformation passes that will often make code worse. However, it’s still worth using, because when your function does autovectorize well, it saves you a ton of work trying to vectorize it manually.