> Are you including the new open source kernel modules?
Apparently, they simply moved almost all of their driver into firmware (the "GPU system processor") - and that firmware is closed. Stated [here](https://github.com/NVIDIA/open-gpu-kernel-modules/issues/19#...).
And what about the host-side library for interacting with the driver? And the Runtime API library? And the JIT compiler library? This seems more like a gimmick than actual adoption of a FOSS strategy.
Just to give an example of why open sourcing those things can be critical: Currently, if you compile a CUDA kernel dynamically, the NVRTC library prepends a boilerplate header. Now, I wouldn't mind much if it were a few lines, but - it's ~150K _lines_ of header! So you write a 4-line kernel, but compile 150K+4 lines... and I can't do anything about it. And note this is not a bug; if you want to remove that header, you may need to re-introduce some parts of it which are CUDA "intrinsics" but which the modified LLVM C++ frontend (which NVIDIA uses) does not know about. With a FOSS library, I _could_ do something about it.
> Out of curiosity, not direct enough for what? What do you need access to that you don’t have at the moment?
I can't even tell how may slots I have left in my CUDA stream (i.e. how many more items I can enqueue).
I can't access the module(s) in the primary context of a CUDA device.
Until CUDA 11.x, I couldn't get the driver handle of an apriori-compiled kernel.
etc.
> Which features are you referring to?
One example: Launching kernels from within other kernels.
> Are you suggesting that features that make programming easier and features that users request must not be added?
If you add a feature which, when used, causes a 10x drop in performance of your kernel, then it's usually simply not worth using, even if it's easy and convenient. We use GPUs for performance first and foremost, after all.