Having different algorithms for different purposes is fine. For instance, autoencoders can do unsupervised learning, GANs can learn generative models that sample from the entire distribution, recurrent neural networks can handle time series, etc. Also while there are many different types of networks, research has shown which ones work and which ones don't. Few people pretrain autoencoders or bother with RBMs anymore, for instance. And I think we have good theoretical reasons why they aren't as good.
But to continue the analogy to computer science, imagine all the different kinds of sorting algorithms. They will each work better or worse based on how the data you are sorting is arranged. If it's already sorted in reverse, that's a lot different than if it's sorted completely randomly, or if it was sorted and then big chunks were randomly rearranged.
There's no way to prove that one sorting algorithm will always do better than another, because there's always special cases where they do do better. The same is true of neural networks, it's impossible to formally prove they will work, because it depends on how real world problems are distributed.