My program performed complicated processing of large 8-bit grayscale images. Naive use of portable SIMD made several important parts of the algorithm run ten times faster than my integrated-GPU implementation! (OpenGL ES 3.0 over ANGLE on an Intel HD Graphics 520)
I think the sheer power of modern SIMD has slipped past most peoples' radar. A multicore processor with AVX2 has number-crunching throughput comparable to a low-end GPU, and for some workloads, the extra flexibility of the CPU can unlock an additional order-of-magnitude performance uplift. (The comparison may be a little different if you're programming the GPU with modern compute shaders - I'm not sure.)
Portable SIMD programming is also just fun! Trying to keep my working set within the 32 KiB L1 cache, trying to keep my data packed within individual 8-bit lanes, and lacking access to some basic operations like integer multiplication and division - it almost feels like programming for a retro games console.