I disagree that his points are due to mere familiarity with "old-school methods". I view elegance as something that is simple in hindsight but only if with an appropriate background. These are simple movements that belie the expertise it took to get there.
Older methods were proof heavy and required more mathematical sophistication but once you understood them, you could easily implement the idea. A lot of deep learning is comparatively mathematically simpler but there is often much incidental detail and folk knowledge that ends up being important but not mentioned. It is in this sense that one can call DL inelegant. A lot of that is because so many are racing to publish results, there's not much time to contrast with prior art or do much more than justify with often very expensive experiments with hastily scrawled descriptions. Lots of ideas generated without much context means many of them are quickly dropped and forgotten, sometimes without sufficient justification.
> Examples of elegance in deep learning
GANs, VAEs, gating, convolutions and weight sharing are indeed great ideas. IIRC Hinton's early papers did inspire Friston's influential work in theoretical neuroscience. However, translation invariance is double counting the advantages of convolutions. Policy gradients, the log-derivative and reparametrization tricks are independent of deep learning. Variational inference is more due to how far reaching that idea is if you want efficient generative models. Game theory...well lots of ideas are connected to it, including evolution and the dead simple weighted majority algorithm. Split brain is really reaching, especially since it's now looking to be one of those ideas that will need a decent amount of revision.
That said, I agree that too much emphasis is placed on interpretability. If a function is too complex, then a human will simply be unable to fit it in their working memory. Nonetheless, the ability to introspect on some of a neural net's decision is vital. As you say, visualizations and mappings which project to a simpler function space but preserve most of the detail are just as possible with neural nets as with any other method.
In this post I might come across as disparaging of deep learning, but this is certainly not my intention. There are lots of excellent papers which elegantly tackle difficult questions. They ask: what invariances and structures do we seek? How do we learn good representations we can sample from? How do we keep learning stable and achieve good gradient flow? How can we capture longer range correlations? These also offer insights that are not limited to deep learning. You just don't hear as much about them because they are not as shiny.
And if you're looking for mathematical elegance that cuts to the heart of the matter, specially with the rising importance of generative models, you could hardly do worse than start from anything written by Shun-Ichi Amari.
People always want to jump right to the shiny stuff and skip the basics, this is understandable but suboptimal in the long term, in any field. But a good compromise on this matter is the excellent free deep learning book by Goodfellow, Bengio and Courville.