Understanding the generalization of ‘lottery tickets’ in neural networks
ai.facebook.com
ai.facebook.com
- The core hypothesis is that within "over-parameterized"[0] networks, a small subnetwork (a small set of weights) is often doing most of the work. A weight can be thought of as an edge in the neural network graph, and so subsets of weights can be thought of as subnetworks.
- You can find these subnetworks by initializing weights randomly, training for some number of steps, identifying the least contributing weights, and then retraining from the same initial parameters as before, except with the least-contributing weights from the previous run zero'd out.
- People have observed that you can achieve something like 99% weight pruning with relatively little loss in performance. After that, things get very unstable.
- This has implications for understanding how neural nets do what they do and for shrinking model sizes.
This is all from memory, so forgive any errors.
[0] - Networks have grown to enormous numbers of parameters lately, and there's reason to think that even before the era of 1B+ parameter networks, neural nets were over-parameterized. Why do we use such large networks then? For a given trained neural net, there might be a much smaller one that does the same thing, but in practice, it's difficult to get the same performance by training a smaller network. This may be because our typical optimization methods aren't well suited for finding these lottery tickets.
Also, because it's cheap to run a network with a billion parameters. Lots of real-world applications are constrained not by compute, but by size of the training set.
Audio-visual inputs do sometimes end up having 1B+ individual values, but there isn't necessarily a 1:1 relationship between input size and parameter count. In many deep neural nets, most of the weights are internal. They're used to process the outputs of earlier layers of the network. This is where the term "deep" comes from.
In addition L2 tends to discourage sparsity by spreading out the influence of weights, which seems antithetical to the mission of pruning. (for example, if you run a ridge regression with two identical features the L2 penalty will assign equal coefficients to both instead of zeroing one out like L1 does)
In an autoencoder, you learn a representation for your data. So you'll get a function mapping from, say an image, to a representation vector. Then you can optimize for desirable qualities for this representation vector, such as sparsity. This will hopefully result in meaningful axes in the representation space, which kind of indicate the presence or absence of different aspects in the input (e.g. whether the input contains a face with glasses or without, etc.). This is an unsupervised approach and results in a (typically lossy) compressed version of the input.
In network pruning you take a previously trained (typically, but not always, in a supervised manner) network and remove individual parameters (weights) from it (i.e. set them to zero) while trying to preserve as much accuracy as possible. Here the trained network itself is compressed.