Basically everything you said was wrong..
The other problem with simd is that in modern cpu-centric languages it often requires a rewrite for every vector width.
Nope, you can use many existing libraries that present the same interface for all sizes, or write your own(which is what I did).
And for 80% of the cases by the point there is enough vectorizable data for a programmer to look into simd, a gpu can provide 1000%+ of perf AND a certain level of portability.
Transfer latency & bandwidth to GPU is horrible, just utterly horrible. And GPU to CPU perf dif is more like 5x, and in games etc the GPU is already nearly maxed out
And it also takes a lot of space on your cpu die. Like, A LOT.
Relative to the performance it can offer it is a very small area. The gains in the article are small compared to what I see, probably the author is new to SIMD.*