CUDA does things like depend on specific scheduler behavior for GPU threads in order to guarantee forward progress, allowing more efficient single-pass computation routines. Or allowing CUDA kernels to launch additional kernels or allocate GPU memory from within the kernel, without host communication. Or checking the GPU architecture and performing specific micro-optimizations designed around that hardware's internal design.
GPU's are nothing like CPU's, in that the x86 architecture is fairly stable. There's not a _ton_ of difference between different CPUs. Sure, performance characteristics and cache sizes might be different, but generally they have the same instruction set. Each GPU generation is basically a completely different architecture, much less between GPU companies. That's (partly) why Vulkan is so complicated - it tries to support the lowest spec 2013 mobile GPU, and the highest spec 2023 desktop GPU in one single API.
The corresponding GPU software complexity is something that feels not quite appreciated. Yet if we are going to see the massive adoption of GPU's that the market projects, with multiple providers of hardware, we'll need some serious advances of the software side.
Being able to describe GPU computations in a declarative way, with a driver similar to an SQL query optimizer, distributing the load.
https://community.amd.com/t5/rocm/available-now-new-hip-sdk-...