This may be a silly question but - so rather than train a big network and hope a subnetwork wins the lottery - why not just train a smaller network with multiple runs with different starting weights?
In this case a positive benefit of combinatorial complexity.