Maybe the additional parameters give the entire network more "leeway" to find such subnetwork structures, i.e. ease gradient descent by smoothing the loss landscape?
That's intuitive but doesn't support the result of the lucky subnetwork (once found and re-initialized) training faster and outperforming the original.
This does not seem to be a contradiction. Once you are in the right region of solution space training is expected to be faster and easier. Re-initialization could have a regularizing effect, explaining the better performance.
They re-use the same initialization, so it appears that the initial weights are inherently coupled with the nonzero structure.