CUDA 11.0
docs.nvidia.com
docs.nvidia.com
This is generally indicative of how poorly organized the CUDA documentation and installation instructions are. The Conda dependency manager has made this a lot easier recently. Especially by, e.g., providing pytorch binaries. Though if you want to use packages like NVIDIA Apex for mixed precision DL[0] you're going to be in for a huge headache trying to compile torch from source while also managing your cuda and nvcc version, which sometimes must be the same but sometimes can not be![1]
[0] Yes, I'm aware that Apex was very recently brought into torch but it seems that the performance issues haven't been ironed out yet.
[1] https://stackoverflow.com/questions/53422407/different-cuda-...
https://forums.developer.nvidia.com/t/the-cuda-toolkit-v10-0...
> The Conda dependency manager has made this a lot easier
Yeah but conda is "Let's do dependency management with a SAT solver, it'll be great!" On a good day, it's just slow. On a bad day, the SAT solver spins for hours before failing to converge. On a really bad day, the SAT solver does something "clever."
I've had a couple of really bad days this year. I'm really starting to not like conda very much.
This means that keeping a conda installation up to date is often very tricky, when upgrading you frequently have to uninstall and reinstall some packages.
It works better if you start from scratch with a requirements.yml file.
Debian managed something like this over 20 years ago in dpkg. But somehow people must keep reinventing the wheel.
Images are publicly available here in case anyone else needs something similar: https://hub.docker.com/u/uodcvip
Furthermore it seems like even the CUDA runtime is typically not installed in the container, but rather injected in by the nvidia-docker container runtime.
It is not fun to deal with.
Here's an example Dockerfile: https://github.com/dmm/docker-debian-cuda/blob/master/Docker...
And here's an example docker run command:
docker run -it --rm $(ls /dev/nvidia* | xargs -I{} echo '--device={}') $(ls /usr/lib/x86_64-linux-gnu/{libcuda,libnvidia}* | xargs -I{} echo '-v {}:{}:ro') dmattli/debian-cuda:10.0-buster-debug /bin/bash
Verbose but it works fine. You still have to have the nvidia driver installed on the host system.
Last time I tried it the cuda inside the container tough it was using some old driver version while a much newer version was installed on the host. So I had to manual install the older version, not sure where the issue was but maybe it was because I was using the deprecated nvidia-docker version 2 which is still needed to pass gpu resources to containers run inside kubernetes.
Something good about Singularity (which I bet you could also do with Docker) is that it automatically binds the right NVIDIA stuff into the container. It also works fine unprivileged :)
This is the kind of thing that happens when you're dealing with a monopoly.
Like Torvalds says [1]: Fuck You, Nvidia.
the Linus video is awesome though :-) And I totally understand his sentiment
docker and nvidia-docker work fine for me
This exact sentence is listed both under "New Feature" and "Known Issues". I'm not super familiar with CUDA stuff, but, it can't be both right?
Even with OSS projects discussions about ending support are not easy.
Also, GCC 9.x compatibility may seem minor to some, but is significant for others. I also think there's some C++17 support in kernels - that's something too.
Edit: oh wait I think I see. Latest supported gcc for CUDA 11 is gcc 9.x, but I think latest Fedora is on gcc 10.
https://docs.nvidia.com/cuda/cuda-installation-guide-linux/i...
Huh? I've been using CUDA for a while now on my Ubuntu 20.04 machine
Just today I've already spent 30 minutes trying to start x with this latest cuda update. Too bad I can't switch back to the open source nouveau driver.
Otherwise, get a pre-owned GTX950 (one that doesn't require external power supply) and a TB3 to PCI-E x16 adapter. Not enclousure, adapter. Should cost you around $200 all in IIRC. And it allows you to upgrade the card furthur down the line since most of the cost is the adapter.
https://developer.nvidia.com/embedded/jetson-nano-developer-...
Here[4] is a notebook that shows how to install CUDA into an environment using the GPU accelerated runtime.
Only major downside is that resources aren't guaranteed (see first section under "Resource Limits" here[5]), so you sporadically may not be able to start a GPU-accelerated runtime session. But that shouldn't be much of a blocker for tinkering purposes.
[1] https://colab.research.google.com/notebooks/intro.ipynb
[2] https://colab.research.google.com/notebooks/gpu.ipynb
[3] https://colab.research.google.com/notebooks/tpu.ipynb
[4] https://colab.research.google.com/github/ShimaaElabd/CUDA-GP...
Its quite easy to set up as well, basically a workstation that you can just connect to with remote desktop, but migrate the hardware it runs on.
The fact that Apple is trying to kill OpenGL and OpenCL and block Vulkan definitely sucks though for anyone trying to do indie games, or open source ML/HPC.