Nvidia-Docker: Build and Run Docker Containers Leveraging Nvidia GPUs
github.com
github.com
(Disclaimer: I work at Docker, but not on the core team)
I'm curious why volume mounting is not an option here. Seems like you'd just ship driverless containers, forcing the user to mount the host drivers before allowing the container to run.
We provide more details here: https://github.com/NVIDIA/nvidia-docker/wiki/NVIDIA-driver Hope that helps!
Yes, we have been discussing with a few people at Docker, I hope attending DockerCon and discussing with the team will help us move forward.
Let us know if there's anything more we can to help.
We haven't found this to be particularly brittle or problematic in any way. It just works(TM). it feels like reaching deeper into the control stack by 'wrapping' Docker CLI or other actions is likely to be an untenable workflow for us.
We use a huge amount of GPU. We do it on CoreOS/k8s et al. We're exploring rkt. Doesn't nvidia's approach interfere with this?
However, I am always looking for better approaches to help reach as far into the future with our platform as possible. So show me the magic, and I'm yours.
1: https://github.com/Avalanche-io/coreos-nvidia (Old public repo, Should I push our latest? I don't think anyone is using this so I haven't.)
What you are showing[1] is how to install NVIDIA drivers on CoreOS the hackish way (not persistent, no driver libs, no DKMS, no UVM, no KMS...)
Regarding rkt, it's not supported at the moment but a similar approach could be taken. As for the Docker CLI wrapper, you can avoid it if you really need to.
I run the driver container at startup, and never shut it down. How is this not persistent? DKMS and other build/deployment choices are not obviated by my approach, so I'm not sure that's relevant.
Looking more deeply at the "Why NVIDIA Docker" in the repo wiki doesn't provide any enlightenment either. In fact it doesn't really explain why docker itself must be modified. The only explanation really is lack of container portability, but driver containers are portable within the scope of a given kernel version. Certainly modified docker cli and plugin requirements are much less portable.
It seems to me like someone at nVidia simply didn't realize that they could run a container in privileged mode and effectively install the driver system wide for all containers.
I'm not going to dwell on the details but there are many reason why doing so can go horribly wrong. Believe me, we (NVIDIA) evaluated our options and know the implications of running our drivers within containers.
Do you really know what --privileged do? If so, you know that there is no such thing as "install the driver system wide". For that you would have to circumvent the namespaces and a bunch of other things that Docker put in place.
"portable within the scope of a given kernel" [and driver] "version"
Well that's not what I call portable :) With nvidia-docker you can build a CUDA image on your laptop and deploy it anywhere in the cloud or on premises without a single modification.
https://github.com/emergingstack/es-dev-stack (feedback/contributions welcomed)
Having a standardised and streamlined way of deploying the application code with CUDA toolkit would have saved me a lot of troubleshooting.