Smarter Training of Neural Networks
csail.mit.edu
csail.mit.edu
> “It was surprising to see that re-setting a well-performing network would often result in something better,” says Carbin.
This, intuitively, makes sense to me. It seems that the pruned model has to waste less training cycles on inferior weights, and it can therefore spend more cycles on further optimizing the good weights.
- "The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks"
Reviewer 2, 3: 9/10 great.
Reviewer 1: "The paper seems a bit preliminary and unfinished."
Authors go on to win best paper with their submission.
Another interesting point is that NNs can be thought about as preforming coordinate transformations on the data manifold, which means these sub nets are potentially approximations to those transformations, potentially up to some scaling factor.
I'm excited to see where this goes
Why?
There may be some in here
They train block-sparse neural networks, where sparseness learned during training.
In the article, if I understood correctly, what they propose is train your network once, remove the x% of smaller absolute magnitude weights, retrain your network fixing those smaller magnitude weights to 0 starting from the same initial starting point.
The idea behind is that your optimization process the first time is telling you that the solution is near a subspace where the weights are 0 but it can't really converge to it. So you project to this subspace by enforcing the weights to 0. Then you retrain again and the search will be easier because the space is smaller, but because you are starting from the same starting point you are kind of guaranteed that you will be able to reach the same optimum but projected.
The problem of the sparsity in the article is that while some weights are 0, they are 0 through masking, therefore you are still doing the computations, and you don't really benefit from the sparsity. If you have enough 0, you can benefit from the sparsity by using some sparse representation, but those are typically an order of magnitude slower than the dense representation.
Combining the idea of the article with the idea from OpenAI of block-sparse neural networks which reduce the operations done without suffering too much from the non-locality and indirection of a sparse representation.
After training normally (provided you have enough memory) (eventually with a sparse-block regularization term to help induce block-sparsity) you may try to prune in such a way that the least significant sparse-blocks are pruned, therefore you may expect both the boost in speed, and the better accuracy and convergence properties.
This is a kind of two phase search, first we look for a finer structure, then we restart to find the best weights for this finer structure.
And they are not "order of magnitude slower": https://stackoverflow.com/questions/24756534/in-what-situati... Unless you operate in binary, of course.
Results from above suggest that you can leave 12.5% of your weights to be non-zero and get nice 4x speed up (1/8 operations done twice as slow). Accuracy start to drop at about 12.5% of weights remaining in both OP and OpenAI papers.
This means you do not need any training to decide what has to be zeroed.
Concerning the sparsity speed-up, I've tried the tensorflow sparse representation a while back, and it was kind of a high effort, low reward process. You had to drop like 90% of the weights (was reducing accuracy), change your ops (dense, convolutions,...) and use big enough layer size to get a feel you were getting something speed-wise.
The openai block-sparse kernels seems promising I'll give them a try.
The openai paper is introducing operations which are a fast middle ground between dense and sparse operations. You still have to specify the sparsity structure you like. (Although often some random sparsity structure work well).
The MIT paper describe one way to choose a sparsity structure and starting point which will work well in the general case.
There are obviously more available sparse solutions if the block sparsity constraint is relaxed therefore I wouldn't be surprised if the best results come from such a network.
You may as well learn block-sparse architecture with 1x1 blocks, effectively doing what MIT was doing, but without two phases.
The idea is that you black out a set of neurons/filters and then train for a short while to overcome the performance penalty. To find the "set" of blacked out cells you could use a genetic algorithm or something, gradually increasing the number of masks.
The last step would be rearranging the network such that the not-blacked-out cells are contiguous, but form smaller layers.
And I remember Hinton hinted at replacing multiple layers by one layer (or no layer), or big layers by smaller layers through retraining.
[1] https://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.32....
[2] https://papers.nips.cc/paper/647-second-order-derivatives-fo...
Such a toxic atmosphere, I don't know why I bother posting.
I am well aware of what meta learning is. If you think that because the idea above comes under the umbrella of metalearning it isn't an interesting idea, then I don't know what to say to you... and the fact that it's downvoted just goes to show the lack of creativity and intuition exhibited by your average hackernews member.
This practice is well known but here's a concrete source from Andrej Karpathy:
https://karpathy.github.io/2019/04/25/recipe/
"init well. Initialize the final layer weights correctly. E.g. if you are regressing some values that have a mean of 50 then initialize the final bias to 50. If you have an imbalanced dataset of a ratio 1:10 of positives:negatives, set the bias on your logits such that your network predicts probability of 0.1 at initialization. Setting these correctly will speed up convergence and eliminate “hockey stick” loss curves where in the first few iteration your network is basically just learning the bias."
Yes, neural networks are objects that compute; there's even this "universal approximator" theorem that says a basic, albeit sufficiently large, neural network can approximate any arbitrary function (from a broad class of functions) to arbitrary precision. However, the theorem says nothing about whether you'll ever actually _find_ the neural network that corresponds to that function. This is what training is for, it allows us to find (the parameters of) the NN that we want to do some computation.
In other words, training is how we program NNs, but in general it can be really hard to arrive at the "program" you're looking for.
Indeed! Most of the time computers are computing, not being programmed. Yet most of the time, neural networks are being trained instead of computing.
That was exactly my point.
But also note that what you say is not necessarily true that the NN's spend most of their time training. Maybe you've got to spend a week on a huge GPU cluster training some autonomous-driving algorithm, but then it runs in "compute" mode for hours a day in tens of thousands of cars.