Having a lot of math involved does not mean that it is mathematically elegant. As a mathematician, I ask myself several questions. First of all, what is really a neural network? Is it an approximating function? Is it a geometric separation on a space, such as SVM? Is it a manifold classificator? (see
http://colah.github.io/posts/2014-03-NN-Manifolds-Topology/, which is very interesting)
Also, what are we approximating? Continuous functions? Non-continuous functions? Are they even functions and not probability measures? Are those functions arbitrary or do they represent something like a manifold?
And the most important: how well are we approximating whatever we want to approximate? The universal approximation theorem gives uniform convergence for measurable functions, but do not specify at which rate or depending on which parameters. It is a strong theorem but not that surprising from the mathematical standpoint, where you already know that you can approximate any function by continuous, compactly supported functions.
Finally, how do you mathematically define the problems that arise in neural networks? What is overfitting? How does the learning algorithm affect the results?
The fact that some techniques are justified by mathematical explanations does not mean that it is mathematically elegant. For it to be mathematically elegant you should have at least clear definitions of the objects of study and the problems you want to solve. I don't think this is the case in neural networks.