This is fair, and if you've got the time and inclination, I'd love to hear about your experience and the tricks you ended up pulling. There are definitely advanced areas of CUDA, and you can go deeper on nearly anything.
But we are in a comment chain spawned by:
> CUDA is fairly straightforward for many tasks and in many cases there is an easy 100x improvement in processing speed just sitting there to be had with <100 lines of code.
And a follow up comment about how easy it would be to write that "<100 lines of code", so I feel like we're definitely talking about the easy case of naturally parallel calculations, and sticking to that as an intro seems fair to me.