The easiest intro IMO is check out & play around in ShaderToy. You don't have to know much about SIMD to write shaders, but once you start paying attention to how the machine works, you can make it go really fast.
In general, using something like CUDA is similar to C++ you just have to make sure all your threads do as close to the same thing as possible in order to see good perf.
It's incredible useful if you're doing a lot of stuff involving matrices (graphics, image/video processing, neural-net stuff, stuff like that). If you have any interest in that topic, it's absolutely worth learning how to use SIMD (or at least learning a library that takes advantage of SIMD in your language of choice).
There are a number of structured approaches to rewriting code to be SIMD friendly, but you do need a degree of explicitness to get what you want out of the compiler.
Vectorization is not always faster. It's important to understand that modern processors can perform work on >100 instructions in a given cycle, and not all instructions take equal amounts of time. So reducing a dozen instructions to a single instruction doesn't necessarily mean that the single instruction is going to be faster.
A single core can have >100 instructions in flight, but most of them will be waiting for their operands to arrive. The theoretical maximum throughput is gated by the instruction decoders and execution units, and is an order of magnitude lower.
I think that we also miss a piece of the puzzle in terms of what happens once you leave ISA instructions and get into uops. There's magic inside there and it's completely opaque.
Are you saying dozen scalar instructions can be faster than one vector instruction? That's wrong 100% of time on modern CPUs.
Some low-performance ARMv7/8 designs can split NEON SIMD instructions into multiple clock cycles, but I think even then NEON is going to perform better.
SIMD is highly optimized at this point, with sustainable two instructions per clock throughput. It's hard to imagine scalar getting anywhere near even at 4 inst/clk rates.
There exist some vector instructions that are going to be slower than a non-equivalent sequence of scalar instructions: VPGATHER is going to be an easy such case.
However, I doubt there are going to be any cases where a vector instruction will take fewer clock cycles than its equivalent scalarized instructions. There are some where it might be equivalent--a vector of 2 elements performing an operation that can be issued twice a cycle is an easy example--but I can't think of any where it would be worse. If that were the case, then you should just implement the operation in hardware by scalarizing the uops (and some instructions appear to be so implemented--e.g., gather/scatter).
On Haswell where it was first introduced... yeah, not very fast, like you mentioned.