Wait, looking at that link I don't see how it avoids downloading CUDA or ROCM. Do you use MLIR to compile to GPU without using the vendor provided tooling at all?
We do use ROCm and CUDA. Only we sandbox it with the model and download only the needed parts which are about 1/10th of the size.