My experience is different. I had to code various instances of "parallel reduction" and "prefix sum" and it's not easy to get into it (took me a day or two). Moreover, coming from an age where 640KB of RAM where considered quite enough, truly realizing the power of the GPU was not easy because my tasks are not quite parallel and not quite thread coherent (I grant you that doing naturally parallel stuff is dead simple). It took me a while and a lot of nvidia-nsight to max out the GPU... Moreover, I was a bit slow to actually understand how powerful a GPU is (for example, my GPU gave poor performances unless I gave it a problem big enough, so I was (wrongly) disappointed when testing toy problems)
But once that challenge is overcome, GPU truly rocks.
Finally debugging complex shaders (I do some specific case of computational fluid dynamics where equations are not that easy, full of "if/then" edge cases, etc) is not fun at all, tooling is sorely missed (unless I've missed something)