Understanding Deep Learning Through Neuron Deletion
deepmind.com
deepmind.com
http://journals.plos.org/ploscompbiol/article?id=10.1371/jou...
And the main part of the paper is that destructive tests in large systems have weird cascading failures that don't generally point to the purpose of the component that failed. Combining that with a domain where you don't have great feedback into the veracity of your conclusions leads to specious results at best.
yep, that was my whole thesis dissertation. also you can actually include some confusion during the training itself by adding noise to the weights or the transfer function and the learning algorithm will find more stable solutions that are more robust on its own.
we used that to produce network that were more resilient to being quantized when moved to low power devices having only fixed point arithmetic available.
it was 2004-2006, we used weight quantization as noise function and a modified levemberg-marquad for training, matlab was the hot stuff back then and by today standard everything is ancient, so there's that - the training alghorithm was changed to train toward a flatter area of the noise function, so it used the second derivative to hold gradient descent and overfitting.
I can answer specific questions if you want.
In dropout, you select 50% (say) of the nodes in one layer of your network at random and force them to pass on no signal (neither positive nor negative). You then train the network to still return the result.
This is effectively training the network to be robust to failures (i.e. "harder to break").
While writing this, I've realized that we've both assumed that not only "networks which generalize better are harder to break" but also "networks which are harder to break generalize better". I think that's true, but worth being explicit.
(Still developing my intuition for machine learning, so I could be wrong...)
So, for example, if you train the network with a 50% dropout rate, dropout will encourage the network to be robust to dropping 50% of the units, but the network could completely fail once 51% of the units are deleted and the training objective would be perfectly happy. As a result, dropout doesn't change the shapes of ablation curves, but rather simply horizontally scales them such that the left edge of the curves is at the dropout fraction rather than 0. In contrast, we found that batch normalization actually pulls the curves up and to the right, rather than simply scaling them, though we only have hints as to why that is.
Hope that was helpful!
also the network could have weird results for dropout values around 20%-30% depending on how the robustness was 'learned'
"So the fact that they're saying they have the same effect as more random ones is interesting". Yes, they are saying that, but their graph doesn't show that. I think their graph supports your intuition, that if you ever plan to delete neurons from a NN, while minimizing the degradation, you should start with the low entropy neurons first.
This may simply reflect the fact that "selective" neurons are relatively rare. The chart in the article implies that there are equal numbers, but it's not clear whether the X-axis is linear or not, or where the threshold is for a selective neuron vs a non-selective one. If 95% of all neurons are non-selective, then the statement above is not surprising.
Dropout is done _during training_ of the neural network. In other words, for a given training example / batch of examples, some of the neurons are deactivated. Those that remain are forced to "pick up the slack" of the missing ones, and so this ends up preventing over-fitting.
What this is, is taking a neural network which has _already been trained_ and seeing how the output differs if we now drop some of the neurons.
It is true that neural networks that have been trained with dropout will probably fare much better under these circumstances.
Would be interesting to have 'resilency' metric in addition to the normal 'accuracy' and 'loss' metrics.
Also wow! Yann LeCunn has authored it
ah ! i was kind of hard pressed for time when i posted that, with the ulterior motive that some kind soul would post the link :)
thank you kindly !
More training data, make the size of the model as small as possible - something I have done since the late 1980s. Not only are neural network engineering techniques rapidly improving, but so is our intuition into how they work.
I think this is key why (generalization) humans only need 1 or 2 instances of something to learn it, while deep learning requires thousands of samples to learn something.
Seems analogous to our learning, where we initially give the subject all of our attention, then infer patterns and use them.