The article meant elegant in the sense of "nn training provably converges in subexponential time and doesnt overfit the data". E.g., like the theory for convex SVMs.
Mathematically, nn training is an unproven algorithm, it just works empirically surprisingly well. It's analogous to the simplex method which worked well for years without theoretical justification.