Of course we need to keep the gradients, architecture search, hyperparameter tuning, scalable training to massive datasets, etc., but there is a growing sense that the programs we write can encode extremely powerful priors about how the world works, and to not encode those priors leaves our learning algorithms subject to attacks, bugs, poor sample efficiency, bad generalization, weak transfer. Not to mention a host of rickety conclusions that are probably poisoned by hyperparameter hacking.
Conversely, we need to try to avoid the proliferation of black box systems that require heroic efforts of mathematical analysis to understand and debug. Take for example the highly sophisticated activation atlas work by Shan Carter and others, which was needed to reveal that many convnets are vulnerable to an almost childlike kind of juxtapositional reasoning (snorkler + fire engine = scuba diver). Beautiful work, but to me it would be better if that form of analysis wasn't necessary in the first place, because the nets themselves were incapable of reasoning about object identity using distant context.
We need systems that are, by design, amenable to rigorous and lucid scientific analysis, that are debuggable, that admit simple causal models of their behavior, that are provably safe in various contexts, that can be straightforwardly interrogated to explain their poor or good performance, that suggest modification and elaboration and improvement other than adding more neurons. We need to speed the maturation of modern deep learning out of the alchemical phase into something more like aeronautical engineering.
The major innovations in recent years have been along these lines, of course. Attention is a great example, basically supplanting RNNs for a lot of sequence modelling. Convolutions themselves are probably the ur-example, of course. Graph convolutions will be the next major tool to be pushed into wider use. To the interested observer the stream of innovations seems not to end. But the framing that makes this all very natural is precisely that of this being the union of computer programming, where coming up with new algorithms for bespoke tasks is commonplace, with automatic differentiation, which allows those algorithm to learn.
What remains exciting virgin territory is how we best we put these new beasts into the harness of reliable AI engineering. That is in its infancy, because how you write and debug a learning program is completely different to the ordinary sort... there are probably 10x and 100x productivity gains to be realized there from relatively simple ideas.