For example, Goodfellow and Bengio's book on Deep Learning talks a lot about connections to Bayesian Inference, Classical Statistics, Information Theory, and a bit about Topology.
Christopher Oolah's blog gives plenty of great explanations of mathematical topics.
Personally, I use Math in my work all the time. Thinking in terms of Information Theory has let me quantify and compare algorithms that seemed difficult to evaluate at first. And several times, I've seen a business make the wrong decision due to lack of Math/Stats background.
I don't mean to be a Math snob – you can absolutely do a lot of valuable work treating algorithms as black boxes. But the Math is there, and it is being used.
https://en.m.wikipedia.org/wiki/Vapnik–Chervonenkis_dimensio...
Also, the reason why black box methods are such a big deal now is precisely because controlled/engineered methods turned out to be inferior (obvious example: image recognition; look no further than into the story of the dropout method and Alex Krizhevsky).
Edit: s/basically forgotten/often forgotten/ in the first sentence.
[0] http://www.math.ucla.edu/~chernikov/teaching/Combinatorics28...
I.e. in many materials or chemistry fields where you start operating with over 100 variables, you begin to develop a dark arts of understanding because what is being attempted is beyond the ability of computers to model.
You can do chemistry without a PhD, but your ability to systematically try to address the complexity may be hindered without the training. Likewise with ‘ML experts’ (I hope they get a cool word some day to describe their profession).
If we have a function f(x1, x2, ...) of 100 variables and you wonder about the gradient, theres multiple ways of calculating it.
Theres symbolic differentiation, but due to the chain rule, the number of terms grows rapidly and the expression can not be stored.
Then theres the finite difference method, whereby you calculate for each of 100 variables x_i:
f(x1, x2, ..., (x_i +epsilon), ..., x100) - f(x1, ..., x100)
the term on the right is the same constant so in total you need 100+1 forward function evaluations. And theres the issue of precision for small differences (mantissa).
One of the main reasons machine learning took off is because of the mathematical realization (Automatic/Algorithmic Differentiation) on how a 1 forward and 1 backward pass is more mathematically rigorous (calculates gradient vs finite differences) and much more efficient.
With the blackboxing of the algorithms, many endusers of the ML libraries end up using ML when they don't know the functional form of a map, but will refuse to apply automatic differentiation of a known complex function with large number of parameters. In contrast those endusers that made sure to understand Automatic Differentiation as a tool orthogonal to arbitrary function approximation (i.e. everyone who realizes the math part of ML is very important) will be able to apply AD (or any other tricks learnt through a mathematical perspective) in situatins where there is no need for arbitrary function approximation...
EDIT: woops I thought you were arguing for blackboxing, against mathematical interpretation upvoted
After all if intelligence comes out artificially, I hope it does not have rule.