Is there some sort of intuition here that would connect this paradox to the effectiveness of dropout in deep neural networks, or am I drawing connections where there aren't any?
Is there some sort of intuition here that would connect this paradox to the effectiveness of dropout in deep neural networks, or am I drawing connections where there aren't any?
The original dropout paper's handwavy justification for dropout is that it prevents co-adaptation. It prevents individual units (nodes/neurons) in the network from relying on specific units in the previous layer firing as well. This is a bad thing because it's fragile (if one unit is off). I say handwavy because this is just intuition; there is not really any proof that this is actually what is happening.
Another commonly cited motivation is that dropout is like learning an ensemble of multiple networks.
The only paper I've seen that theoretically analyzes dropout is: https://arxiv.org/pdf/1506.02142v6.pdf, which proves it's equivalent to approximating gaussian processes (this is beyond me).
Thank you :)