23 karma · joined December 26, 2019
So C is too high level for GPU programming, and needed to be extended.
Because the speed advantage is so huge, people went through the pain of learning this new model and redesigning algorithms to better fit it.
I'm curious if this still stands in general today, given today's advanced branch predictors and the fact that the CPU tends to be memory-bandwidth-bound, thus you have more "free computation" while waiting for memory (what GPU programmers call compute to memory ratio).
This way you can have the speed and memory tightness of C++ where it matters, but do the other stuff, like initial data parsing and munging or overall control and sequencing in Python, thus avoiding the general clumsiness and unproductivity of C++