Deep Learning Without Poor Local Minima
arxiv.org
arxiv.org
[1] http://stats.stackexchange.com/questions/203288/understandin...
It has been know for a very long time that simple models, like the Random Energy Model, displays a spin glass transition at low temperature. (the REM is a p-infinite limit the mean field p-spherical spin glass used in the earlier papers by LeCun) So it is expected that a random network may also display this kind of behavior, and, therefore, there could exist a large number of global minima, separated by very large barriers.
see, for example http://guava.physics.uiuc.edu/~nigel/courses/563/essays2000/...
However, it has been argued that this is unphysical for very strongly correlated systems like proteins and random hetero-polymers (with non-local contacts). Instead, one would have a spin glass minimal frustration. This would lead to a highly funneled energy landscape, with a single (or a few local) minima. This would lead to a rugged convexity.
https://charlesmartin14.wordpress.com/2015/03/25/why-does-de...
It would be helpful if some practitioners could comment on the reliability of the assumptions in this paper.
- That the dimensionality of the output of the network is smaller than that of the input. That is usually the case in image recognition, where the image is width x height x channels dimensional, while the output is usually a much smaller number of label-wise probabilities. It probably isn't the case when you generate data from some smaller representation, e.g. with autoencoders, image generation, etc.
- That the input data is decorrelated, and that the input data is uncorrelated with the output ground truth. The former can easily be obtained via a whitening transformation in many cases in practice. I am not quite sure about the latter.
- That whether a connection in the network is activated or not is random with the same probability of success across the network. Active means the ReLU activation function has output greater than 0. Many people initialize weights in the network with some 0-mean random variable and some constant bias, in which case that assumption should hold true at the beginning of training. Whether that assumption holds throughout training could easily be verified empirically - i.e. by monitoring the network's activation.
- That the network activations are independent of the input, the weights and each other. That's obviously not completely true - the network activations of a given layer are a function of the previous layer's activations and weights, and ultimately of the input in the first layer. With large enough networks, this may hold sufficiently in practice - any single activation should not depend very significantly on any other single variable.
I may have missed something in interpreting the maths, any comment is appreciated. From a practical standpoint, especially for computer vision, these assumptions seem quite reasonable. I am not qualified however to comment on the proof of this result, so I would wait on peer review. Still, it is heartening to see the theory of deep learning finally catching up with practice!
I haven't read it yet. Is that just linear correlation, or full independence? If it's independence, then there's no signal, right?
All this talk of maxima and minima reminds me of Morse theory, for instance (and that Wikipedia page is more than what I know about it). [1]
Is there any sense in what I'm saying?
Quoted:
1. Our labeled datasets were thousands of times too small.
2. Our computers were millions of times too slow.
3. We initialized the weights in a stupid way.
4. We used the wrong type of non-linearity.
The pre-training helped with initialization, but later it turned out that just initializing the weights with correct scales for each layer (to deal with dissapearing/exploding gradient effect) worked almost as well with enough data.Dropout is a tool to prevent overfitting (large gap between training error and validation error). This paper does not say anything with overfitting or generalization or the impact of regularization.
Also note that unsupervised pre-training via stacked RBMs has proven mostly useless for MLPs with ReLU activations if the network is wide enough and the number of samples in the training set big enough. It is unclear that initialization via unsupervised pre-training can improve the training error or not. I think it mostly has an impact on the validation error (although I am not sure).
Furthermore, practitioners tend to stop training before the full convergence on the training set. Instead one generally stops when validation error stops decreasing significantly (early stopping) and one does not really care about the final value of the training loss one could have reached if we had continued training forever. Traditional SGD has a convergence rate that is too slow and in practice it prevents checking whether we are converging to a bad local minima or not on non-toy problems.
To sum up: better understanding of the optimization problem is very helpful (in particular to tackle underfitting and reduce training times) but that alone will not ensure that we can build model that generalize correctly to unseen data.
> For deeper networks, Corollary 2.4 states that there exist “bad” saddle points in the sense that the Hessian at the point has no negative eigenvalue.
To me these sound just as bad as local minima. Also I don't think it's standard to call something a saddle point unless the Hessian has negative as well as positive eigenvalues. Otherwise there's no "saddle", more something like a valley or plateau.
They claim that these can be escaped with some peturbation:
> From the proof of Theorem 2.3, we see that some perturbation is sufficient to escape such bad saddle points.
I haven't read through the (long!) proof in detail but it doesn't seem obvious to me why these would be any easier to escape via peturbation than a local minimum would be, and I think this could use some extra explanation as it seems like an important point for the result to be useful. Did anyone figure this bit out?
I was thinking of examples like (x-y)^2 at zero, although I guess that's still a local minimum, just not a unique local minimum in any neighbourhood.
2) every local minimum is a global minimum
3) every critical point that is not a global minimum is a saddle point, and
Are these not the same thing?This satisfies 2) but not 3).
The point x = 0 is a critical point but it is not a global minimum and it is not a saddle point.
Consider the following question - how many layers do you need to build a RELU network to approximate x^2 function on [0,1] with 1e-6 accuracy.
https://github.com/ghostFaceKillah/tensorflow-experiments/bl...