The size of CUDA really is astonishing. Any chance someone might figure out how to slim that down?
Nvidia is the only one who could, since they own it.
Talking directly to the kernel / driver / firmware.
As others have said, George Hotz is doing his best in reverse-engineering and skipping layers.
Taking a peek inside the package it seems to mostly be the libraries - CuFFT alone is about 350MB for example, twice over for the debug and release versions. I'm guessing those are probably fat binaries pre-compiled for every generation of Nvidia hardware rather than just the PTX bytecode, which would help to speed up fresh builds, at the expense of being huge.
I don't think it's about the byte size, but the inherent complexity of the implementation. 1000 lines of C code is extremely simple by any standard. Whereas a sundry collection of Python and PyTorch libraries is anything but.
A bunch of install methods for torch via pip include ~1.5GB of lib/ because of CUDA. libtorch_cuda.so is like 800MB on its own
I mean, being fair, the 2.4GB CUDA SDK is absolutely required for the cPython implementation as well