CUDA/etc is certainly a pain.
A good number of years ago I wrote my own Torch-like C++ NN framework using CUDA (cuDNN, cuBLAS) for the GPU, and one of the annoyances is that cuDNN has incomplete coverage of even the basic operators needed. Add/Sub/Min/Max/Sqrt/Negate are all provided as part of cuDNN, but if you want other common NN building blocks like Div/Exp/Log/Pow/Inv/InvSqrt then you have to write them yourself in CUDA and either forgo cuDNNs tensor-descriptor layout flexibility or re-implement that yourself too.
Of course frameworks like PyTorch support all the operators you'd expect, since they've written their own kernels in CUDA where the functionality is missing from cuDNN.
It's not clear exactly what Keller is referring to there. When people say CUDA they might be referring to the entire ecosystem (nvcc CUDA C/C++ compiler, CUDA API's, higher level cuBLAS, cuDNN, etc), or maybe just the base compiler (which lets you write your own kernels) and API for allocating memory, queueing kernels, etc.
The cuDNN kernels (convolution, etc) are highly optimized, as is cuBLAS (e.g. matmul), and I doubt anyone is going to do better writing these themselves. Does Keller consider using cuDNN as "writing CUDA" ?
As far as I'm aware the higher level, performant, CUDA libraries, as well as specialized components like TensorRT are written in CUDA, although that could mean a combination of C/C++ & PTX pseudo-assembler (ptxas is really a compiler, not an assembler). The alternative would be they they were written in hand optimized SASS assembler which afaik is only available outside of NVDIA via the Open Source CuAssembler.
I believe Mojo support for NVIDA is based on PTX. I'm not sure if that would be really be considered as "CUDA" or not if they are not using nvcc at all.
Most people, outside of framework vendors, would have no reason to use CUDA anyway, since it's just too low level. The only sane use case would be where writing a custom kernel in (e.g.) PyTorch or Mojo doesn't get the performance you want and you write than one kernel in CUDA. The hope would be that the Mojo compiler is good enough that this would not be necessary.