> is there any math behind ML at all?
Yes. There is a lot of research in this area (some people argue it's excessive at the moment). You have the correct intuition that the answers aren't black and white all the time. For example, there are solid reasons to choose relu activations over tanh. Or to build certain types of network architectures for certain tasks. That doesn't mean that you can immediately calculate what would happen if you switch from one activation to another without running your network.
The paper that introduced batch norm, adaptive instance norm, attention heads, or any module used in a network have an extensive discussion of the motivation for their existance, some derivation or proof that they do what you want, and an empirical test to show it helps in practice. The reason some losses allow GANs to converge in certain situations while others don't isn't a complete mystery, there is theory that supports this.
Researchers designing new models are considering weak points in old approaches, identifying why they aren't working correctly, and proposing something new that solves a part of the problem. All of this is done by looking at the math behind all the operations in the network (or at least the parts relevant to a certain question).
That nobody really knows how AI works is one of those myths told by the media. Just because the model weights aren't interprable doesn't mean we don't know why that model works well. It just takes quite a bit of maths knowledge to really understand state of the art models. All that knowledge is also easily packaged into modern frameworks that make it easy to use without a deep knowledge of why it works. All of this contributes to the feeling that nobody really knows what's going on, while in reality it's onky the majority of people that don't know what's going on ;)
Let's take the simplest example: recognizing the grayscale 30x80 pictures with 0-9 digits. IIRC, this is called the MNIST example and can be done by my cat in 1 hour without prior knowledge. Let's choose the probably simplest model: 2400 inputs are fully connected with a 1024 vector that's fully connected with a 10 vector. And let's use relu at both steps. We know that this kinda works and converges quickly. In particular, after T steps we get error E(T) and E(1e6) < 0.03 (a random guess). Can you tell me how T and E will change if we add another layer: 2400->1024->1024->10, using the same relu? Same question, but now we replace relu with tanh: 2400->1024->10.
My impression is that obviously ML is guided by math and people want to have an understanding of why some things converge and others don't. But "in the field" many people just mess around with different set-ups and see what works (especially in deep learning). Maybe theory follows to explain why it worked. I think you're right that a lot of progress in the field is based on intuition and some reasoning (e.g. trying something like an inception network) more than derivations that show that a particular set-up should be successful. I get the impression that most low-level components are pretty well understood, but when they are stacked and combined it gets more complicated.
I would be very curious to see a video of your cat solving MNIST in 1 hour!
It's not a myth. No one really understands how neural networks work. We don't know why a particular model works well. Or why any model works well. For example no one can answer why NNs generalize so well even when they have enough learning capacity to memorize all training examples. We can guess, but we don't know for sure. Most of the proofs you see in papers are there as fillers, so that papers seem more convincing. We rarely can prove anything mathematically about NNs that has any practical value or leads to any breakthroughs in understanding.
If we did really understand how NNs work, then we wouldn't need to do expensive hyperparameter searches - we would have a way to determine the optimal ones given a particular architecture and training data. And we wouldn't need to do expensive architecture searches, yet the best of the latest convnets have been found through NAS (e.g. EfficientNet), and there's very little math involved in the process - it's pretty much just random search.
Funny you mentioned the batchnorm paper - we still don't know why batchnorm is so effective - the paper gave an explanation (covariate shift reduction) which later was shown to be wrong (batchnorm does not reduce it), then several other explanations were suggested (smoother loss surface, easier gradient flow, etc), but we still don't know for sure. Pretty much every good idea in NN field is a result of lots of experimentation, good intuition developed in the process, looking at how a brain does it, and practical constraints. And yes, sometimes we're looking at the equations, and thinking hard, and sometimes we see a better way to do stuff. But usually it starts with empirical tests, and if successful, some math is used in the attempt to explain things. Not the other way around.
NNs are currently at a similar point as where physics was before Newton and before calculus.
I'm more inclined to compare with the era after Newton and Leibniz, but prior to the development of rigorous analysis. If you look at this time period, the analogy fits a bit better IMO -- you have a proliferation of people using calculus techniques to great advantage for solving practical problems, but no real foundations propping the whole thing up (e.g., no definition of a limit, continuity, notions of how to deal with infinite series, etc.).
Or maybe it's as useful as a rigorous mathematical analysis of a brain - again, not very useful, because for us (people who develop AI systems), it would be far more valuable to understand a brain on a circuit level, or an architecture level, rather than on a mathematical theory level. The latter would be interesting, but probably too complex to be useful, while the former would most likely lead to dramatic breakthroughs in terms of performance and capabilities of the AI systems.
So maybe we just need to keep doing what we have been doing in DL field in the last 10 years - trying/revisiting various ideas, scaling them up, and evolving the architectures the same way we've been evolving our computers for the last 100 years, with the hope there will be more clues from neuroscience. I think we just need more ideas like transformers, capsules, or neural Turing machines, and computers that are getting ~20% faster every year.
Also, the step of "shuffle around the ML graph using some intuition" involves gathering that intuition, which usually arises from a great deal of mathematical competence. A 3x3 conv kernel versus a 2x2 one can, for instance, be discussed in terms of Fourier theory and mathematical image processing, but areas with huge built-in theory.
Things like replacing the activation function were initially studied anecdotally. People realized that in some settings one activation function or another would lead. Eventually, there was also theory showing that in large nets of stable configurations, there was serious interaction between the initialization method and the activation function and problems like poor backprop signal propagation were tackled theoretically and practically.
Generally, the mystery comes from the vast parameterization of these DL models. They operate in a space that's very hard to generalize—large, finite spaces. Small finite spaces get treated exhaustively. Infinite spaces get treated asymptotically. Large finite spaces get bounded on either side by those methods.
So yes, there might feel like there's a dearth of theory in DL when it comes to the large scale behavior of a general network. That can be super frustrating. At the same time, people are trying to push through and create more theory every day.
In addition, it came with tricks with how much noise to inject in certain situations. "How much noise do you need is enough to escape?" which is pretty practical
https://joanbruna.github.io/MathsDL-spring18/
On Tanh and ReLus, the paper on SeLUs is pretty intense.
ReLUs help with vanishing gradients to some degree on a larger degree of loss and activation functions, btw.
edit: On your question about a billion data points and predicting behavior, we are getting there.
One of the most interesting and elegant examples , Topological Data Analysis that is based on Topology and utilises tools like persistent homology. They can be applied to image processing, classification, etc.
[Edit] Whether this model works is question. Sure, it can't recognize dogs or cats, but what if the dataset is 0-9 digits? Now it suddenly works, right? And works really well. But what's changed? Can we describe in mathematical terms what makes the 0-9 dataset so special? It'll probably work with A-J letters too, but what about hieroglyphs?
The math I'm looking for would tell that E(T) is the error and on such model and such dataset, E(T)=exp(-T^2)+O(exp(-T^3)), according to such and such theorem; and according to another theorem, if the dataset is isomorphic to that manifold, E(T) can be improved to O(exp(-3T^2)).
So they settle for much smaller targets. Either of understanding how much simpler systems work. Or of trying to understand a little bit the effect of tweaking something in some more complicated model.
Perhaps you should think of these two approaches as analogous to doing simple chemistry (what shape is a sugar molecule? A DNA molecule?) vs trying out drugs (if you eat the bark of this tree, you don't get malaria! Let's refine that stuff). Both can be useful, but they are very far from a unified theory of how your body works.
Naiver-Stokes is much simpler because it operates by itself. Of course turbulence is hard but even there we usually care about its coarse features, we'd be content to throw away almost all the information provided the calculation of the wing's lift works out OK.