I'm not sure dropout has anything to do with local optima or removing greedily optimal paths, since it is random.
The original dropout paper's handwavy justification for dropout is that it prevents co-adaptation. It prevents individual units (nodes/neurons) in the network from relying on specific units in the previous layer firing as well. This is a bad thing because it's fragile (if one unit is off). I say handwavy because this is just intuition; there is not really any proof that this is actually what is happening.
Another commonly cited motivation is that dropout is like learning an ensemble of multiple networks.
The only paper I've seen that theoretically analyzes dropout is: https://arxiv.org/pdf/1506.02142v6.pdf, which proves it's equivalent to approximating gaussian processes (this is beyond me).