GPU Headaches: Notes on Installing CUDA, CuDNN and Tensorflow on Manjaro
leblancfg.com
leblancfg.com
https://lambdalabs.com/lambda-stack-deep-learning-software
We wrote debian packages for every framework, including cuda and cudnn. Using our Debian repository, you can install all these frameworks using apt/aptitude.
When a new version of a framework comes out, we usually have it available in 1-2 weeks.
Installing this stuff is a huge nuisance for us and we have some pretty insane Dockerfiles to handle all the different combinations. I might look into using this for our ML images.
Seriously this is too cool and nice, theres no way you're so cool and nice, such people just don't exist! (really this is awesome, thank you)
Our company used to run a cluster of ~1000 gpu servers for inference and training. We used caffe and torch. Provisioning and software maintenance was a huge hassle.
We realized how painful it is to get a machine set up for deep learning, so we decided to release our debian packages to the public.
We sell computers built for deep learning researchers and figure the more people we can help get started, the better :)
Manjaro is a great choice and it does "just work" the best. All the others work just fine though.
All the ML researchers lusting for bare metal need to decide if they want to be "Gentoo Ricers" - https://funroll-loops.teurasporsaat.org/ - or if they just want to train models and get real work done.
I compile every Tensorflow version (to compile with march=native). It is usually uneventful and takes a couple of minutes of my time.
By the way, a Tesla card also costs > 5000 Euro and then add the cost of a machine with enough fast cores, memory, and disk space and you are north of 25000.
https://blog.perfinion.com/2018/07/tensorflow-cpu-supports-i...
That said, configuring a GPU machine is (still) ridiculously complicated IMO. It's just that I've been doing it for almost 20 years.
Thoughts:
* If I was setting up multiple identical machines, I'd only need to solve the issue once.
* Details matter: Switching from VGA to HDMI connector made a world of difference. Me not noticing that earlier cost me ~2 days of trial-and-error to go to waste.
export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/usr/local/cuda/extras/CUPTI/lib64
Reboot and cross your fingers.
Rebooting will terminate the process tree in which the environment has been modified, making this line a no-op. Please take the time to understand the commands you run before suggesting them to others.Back to your point, Manjaro runs directly on pacman repos, not a layer between it and upstream. It just comes with IMO sane defaults and an installer GUI.
If not, I think it doesn't make any sense to use it unless you are a seasoned user. Rolling release will make running big frameworks like KDE tricky, as things will change often. And unless you have set everything up yourself, you won't know where to look in order to fix things.
I'd recommend you look into NixOS if you want to run a desktop environment. It's radically different (totally functional and declarative, whereas Arch is imperative). All Python and deep learning stuff is neatly packaged and tested. And rollbacks are trivial. There's no state.
You can even run Nix in other distros or macOS.
On the contrary, you get new releases (with bugfixes) much sooner.
P.S. well you need to install docker with Nvidia support, but you can use it in the future for other things. This is also a couple of packages in Pacman (maybe Nvidia docker is in air, I don't remember now).
I should add that as an addendum to the article, though, great catch!
You can more or less do it with SYCL / ComputeCPP, which is great, but it is not trivial either. It does require you to compile TF and may fail to operate some operations.
Full disclosure I have not retried it recently, last time was 6 month aago, and a great amount of work is done on the SYCL / ComputeCPP front.
If graphic card makers would open source their drivers it would come default shipped apt-get,pacman,dnf,yum installable by the Linux distros.