Running things on some accelerator (gpu, etc.) usually involves writing a specific kernel in a language subset, manually copying data and generally long latencies. Unless there is a lot of data it won't be faster.
With AVX in the best case the compiler can just vectorize some loop, speeding it up 5x without any added latency or source code changes.
See libjpeg-turbo, ffmpeg, crypto, hashing in general, ripgrep, simdjson, ...
https://woboq.com/blog/utf-8-processing-using-simd.html
See also: https://www.reddit.com/r/compsci/comments/4cq0ls/when_is_sim...
If simd was obsoleted by GPUs intel & amd wouldn't keep introducing wider simd extensions
Curious to think about how unified memory may change the ratio of flops/memory access when it makes sense to shift job from CPU (better for low number) to GPU (better for high ratio)