It’s not universally good. It’s just that if you take naïve C code that does bulk operations on arrays, and compile it at -O3, you might get a very nice improvement from the auto-vectorizer. With some additional restrict keyword annotations, you can sometimes improve the performance significantly.
Then there are non-standard tricks you can use, like __builtin_assume_aligned().
This is something of an arcane art. Sometimes it just doesn’t work at all, sometimes it’s not obvious how to write your code so that the autovectorizer can work well, and getting the autovectorizer to work requires enabling various code transformation passes that will often make code worse. However, it’s still worth using, because when your function does autovectorize well, it saves you a ton of work trying to vectorize it manually.
These will be used by default (at O2 without vectorization enabled) for single-precision floating point math on x86-64 (where stack-based FPU is deprecated).
Vector instructions you're referring to end with "ps", they're very unlikely to be emitted for typical fp operations.