I'm sure CUDA is great, and if I had more free time and/or better reasons to improve the performance of my code it would probably be great for me. My point was mainly that a few lines of code which may be trivial for one person to write may not be for someone else with different experience. Depending on what the code is being used for even a vast increase in performance may not be worth the extra time it takes to implement it.
[0] https://docs.nvidia.com/cuda/cuda-c-programming-guide/index....
If you're already familiar with one of the languages that the nvidia compiler supports? Not that many. For people familiar with C or C++, it's a couple extra attributes, and a slightly different syntax for launching kernels vs a regular function call. I'm admittedly not experienced with Fortran, which is the other language they support, so I can't speak to that. There's c-style memory allocation functions, which might be annoying to C++ devs, but it's nothing that would confuse them.
Edit: There's also a couple weird magic globals you have access to in a kernel (blockIdx, blockDim, threadIdx), but those are generally covered in the intros.
But once that challenge is overcome, GPU truly rocks.
Finally debugging complex shaders (I do some specific case of computational fluid dynamics where equations are not that easy, full of "if/then" edge cases, etc) is not fun at all, tooling is sorely missed (unless I've missed something)
But we are in a comment chain spawned by:
> CUDA is fairly straightforward for many tasks and in many cases there is an easy 100x improvement in processing speed just sitting there to be had with <100 lines of code.
And a follow up comment about how easy it would be to write that "<100 lines of code", so I feel like we're definitely talking about the easy case of naturally parallel calculations, and sticking to that as an intro seems fair to me.
And there's also value in seeing how other people approached a problem.