+ Normalization is beneficial for learning (per unit zero means and unit variance). It can be batch normalization, layer normalization, or weight normalization (if trained layer for layer and previous layer normalized).
+ Perturbations through stochastic gradient descent, stochastic regularization (dropout) does not destroy the normalized properties for CNNs, but it does so for forward nets.
+ Self-normalizing net uses a mapping g: O -> O that maps mean and variance to the next layer for each observation. Iteratively applying this mapping leads to a fixed point.
+ The activation function to do so is not a sigmoid, ReLU, etc. but a function that is linear for positive x and exponential in x for negative x; the scaled exponential linear unit.
+ Intuitively: for negative net inputs the variance is decreased, for positive net inputs the variance is increased.
+ For very negative values the variance decrease is stronger. For inputs close to zero the variance increase is stronger.
+ For large invariance in one layer, the variance gets decreased more in the next layer, and vice versa.
+ Theorem 2 states that the variance can be bounded from above and hence there are not exploding gradients.
+ Theorem 3 states that the variance can be bounded from below and does not vanish.
+ Stochasticity is introduced by a variant on dropout called alpha dropout. This is a type of dropout that leaves mean and variance invariant.
I think the paper gives a nice view on handling gradients in deep nets.