Deep Learning and Free Software
lwn.net
lwn.net
Most of the questions raised in the thread were basically irrelevant to Leela Zero, which does everything "correct" from the point-of-view of a strict interpretation of free software:
- open freely licensed data set - open freely licensed training code
The only issue relevant for user software freedom, is that the results of the training process can't be reproduced easily.
I was irritated at the thread because 90% of it was theorising about some alternative situation where the data set was not free, or the training code was not free. That is not the situation we have with Leela Zero, we can skip that discussion and focus on what's actually at issue, i.e. the reproduction / verification of the training process.
I suppose no one could just lend you the infrastructure needed to rebuild the model (sponsors, or otherwise supporters of Debian?)
Whether the abstract execution system is implemented in a proprietary or free way is a separate issue. For example these days it's hard to execute x86 instructions without running some sort of proprietary microcode and/or management engine. For comparison, OpenCL has many competing implementations, including the free open source MESA graphics drivers.
In order to reproduce the current best-weights then, one would have to record exactly what versions of Leela Zero and ELF was used to generate which subsets of data, and which subsets were used to create further subsets of the data. I don't think anyone has kept that information around, so I'd guess the current best-weights will not actually be reproducible, ever.
In future, one can imagine other software that could keep track of this information, and then be actually able to reproduce a particular resulting set of weights.
However let's step back a bit. On a high-level, nobody actually cares that the results are not fully deterministic, they only care that it is a faithful representation of what the source code does. This is true both for software determinism and for weights-model determinism. Being deterministic is a (relatively) easy property, which when we achieve it, allows us to verify that the results (binary software, trained weights) don't contain backdoors or other unpleasantness that's not visible in the source code. But the latter is what we "actually" care about.
If we can achieve the latter property without achieving determinism, then we are also mostly satisfied. That would involve being able to examine the model directly and see what it does, and see that it doesn't contain backdoors or other things. I can't even begin to imagine how to achieve this, it is a hard problem and the solution to this, would also solve the criticism of these AI weights/models being opaque and not really contributing much to human knowledge.
Determinism is still useful for other purposes though. If you know exactly how something was produced, you have much greater control and understanding of how to tweak it, which might actually help us with the aforementioned goal of deeply-understanding these weights from a human perspective.
On the framework side, you'd want good ways to ensure that random seeds are well distributed across a cluster of workers (and be able to restore state of a machine goes down). Also need to ensure that workers consume training data in a fixed order. This will slow you down, since you've got to wait for the slow poke every time, before updating weights.
To me, this is the crux of the matter. If the view is taken that software distributed under a "free-software" compatible license is non-free without the ability to obtain all training sets for any models, there's going to be huge difficulty in incorporating ML into free software in many cases. Since many datasets aren't able to be redistributed freely (e.g. licensing, legal, cost), that's an enormous advantage for non-free offerings.
A possible route around this might be 'community curated' datasets, where contributors freely license their data in the same way as we do code. It'd be interesting to see if they come with an analog of the AGPL - ie. models trained on this dataset must be released to the user as source. (This might already exist?)
Whatever difficulties this view presents, this is the only view that makes sense.
It's like saying you can modify and audit Microsoft Windows or other proprietary software because you can write and read the bytes of its object code.
Modification: There have been repeatedly successful case of transfer learning where you take a pretrained deep neural network and train on a new dataset (even very small).
Vulnerability: you dont need the training dataset to test vulnerabilities, you just either very specific input (that you probably gonna have to create yourself because they are so specific that they are not even in the dataset) or using one input that you corrupt for attacking. And as far as I know it doesn't sound harder than making test for regular algorithms.
Edit: yes. Discussion on hn: https://news.ycombinator.com/item?id=17231593
So far Nvidia does excellent technical work with their hardware + their out of the box CUDA toolkit + support in every major deep learning library. The drawback is that they keep everything closed and they charge a huge premium. AMD has great hardware at good prices but the software side is non-existent and the Linux support goes from bad to worse.
Nvidia is the one having absolute control here and they choose to squeeze the market because they can.
Intel's free software release seems to me to be rather in their favour compared with other vendors.
MKL’s fft isn’t a huge improvement on fftw with avx512, but its blas is ~3x as fast as openblas currently. And before these open source projects caught up, it was by far the best.
And I agree, they’ve been better about it lately. I was comparing Intel then and Nvidia now.
For a free avx512 BLAS, use the current release of BLIS. OpenBLAS recently gained skx (but not knl) gemm support, but I don't know how good it is as I don't have the hardware.
When they realized the mistake and came up with SPIR it was too late.
Also SYSCL is no better than CUDA with their community edition.
First, you can load things into your network regardless if you use CPU or GPU. And if needed, people can write GPU code for other architectures/
Second, inference (prediction) is fast. Yeah, it may be not useful for real-time applications on CPU (like: self-driving cars) but for detecting one object it can be fast (see e.g. https://transcranial.github.io/keras-js/#/resnet50).
Third, using things like TensorFlow.js you get GPU acceleration with any GPU card, not only nVidia. It is not nearly as fast, but still faster than Python + CPU. There are real-time demos such as https://experiments.withgoogle.com/collection/ai/move-mirror....
Side note: I just start https://inbrowser.ai/ for tutorials and open source templates for using fully frontend AI.
Since giving away this data is mostly neither possible or desirable legally, nor in the interest of the dataset owner, there is a tradeoff between wanting to learn from data, and preserving privacy of the people generating/described by the data.
That is not to say that science is not trying, like for example this paper:
Is the era of computers + internet = freedom and knowledge-sharing over?
Perhaps I'm over-dramatising but this is not something to be okay with.
At this point in time you just can't do some of the DL stuff without proprietary nvidia tech.
Sure, you can't do it on ONE CPU, but the point is to have a cluster of CPUs. It's not the case that you're forced to use nVidia's proprietary stuff in order to do deep learning.
[0]: https://arxiv.org/abs/1712.06567 [1]: https://towardsdatascience.com/paper-repro-deep-neuroevoluti...