SIMD, SIMT, SMT - parallelism in NVIDIA GPUs
yosefk.com
yosefk.com
I was surprised that the author didn't once use the term CUDA though, they even discuss actual syntax from it, but don't mention the language (extension) once.
I wonder if what's needed is a higher-level representation that can compile to the best access patterns for the given hardware. (And something that can try several access patterns for your problem and choose the most efficient one.) GPU programming is still quite new, so I guess it's bound to show up eventually.
If it can't handle all possible situations, such a tool would be still be useful, even if you end up having to go down to the CUDA/OpenCL level for certain problems that are too difficult to express declaratively.