However I don't like that there is often a strict dichotomy presented between "deep learning" and "statistics". There is a whole world of gray areas and hybrid techniques, which tend to be both more accessible, easier to reason about, and more effective in practice, especially on smaller "tabular" datasets. What about generalized additive models, random forests, gradient boosted trees, etc.?
The author of the document I'm sure is aware of these techniques, and I assume they are left out because they didn't perform well enough to be considered here. But I don't think it does the discourse any favors to promulgate the false dichotomy.
Vanilla deep learning models are statistical models (a la linear regression) and not probabilistic models (a la Gaussian mixture). It is important to maintain the distinction.
But to your point about the dichotomy between deep learning and more "traditional" statistical methods: this confusion in common parlance clearly has negative effects on model-building among engineers. You are right that when people think "deep learning" they think of very specific architectures with very specific features, and don't seem to conceive of the possibility that automatic differentiation techniques mean you can incorporate all sorts of new model components that blur the line between deep learning and older methods. For instance, you could feed the results of a kernel SVM to an ARIMA model in such a way that the whole thing is end-to-end differentiable. In fact, the great benefit of deep learning long-term is (in my opinion) that the ability to build these compositional models means you can bake in that much more inductive bias into the models you build, meaning they can be smaller and more stable in training.
Isn't this just a matter of interpretation of the models? You can interpret linear regression in a Bayesian way and say that the prediction of the linear model is the MAP of the mean, you can also calculate the variance, the l2 norm objective is saying the distribution of errors is normally distributed, l2 regularisation is a normal prior on the coefficients, etc, etc? All the same stuff can be applied to deep learning models.
Maybe I don't understand your distinction between statistical and probabilistic though?
Not really. This is the classic frequentist vs Bayesian debate. In frequentist-land, you are computing point estimates of the model parameters. In Bayesian-land, you are computing distribution estimates of the model parameters. It is true that there is a difference in interpretation of the generative process but the two choices demand fundamentally different models because of the decision about which of the parameters or data are considered "real" and which are considered "generated".
I think a more abstract/general way to put it is: "statistics" is concerned with statistical summary values (i.e. mean-field estimates over measures) while "probability" is concerned more with distributions (i.e., topologies of measures). I'm not sure this is a rigorously correct way to characterize it, but it illustrates the intuition I'm trying to convey.
There are some probability models that are not really statistical models, but there are few or no statistical models that are not also probability models.
Least-squares regression is a probability model. Even if you don't particularly care about the error distribution, you are still estimating a conditional expectation and setting a conditional independence assumption on the residuals. If that's not a probability model, then I don't know what it is!
probability : distributions :: statistics : expected values
Statistics can be summarizes as one thing n -> N. Does ‘little n’ represent ‘big N’. In other words, does the sample generalize to the population. Statistics means something like “description of the state”. It was born out of census samples where larger population samples had to be estimated. “n” could be a handful of fish in a “N” lake. “n” could also be the parameter estimated in a linear regression with the sample of data collected while “N” is the true parameter of the relationship if we had all the data. Point estimation is about finding the needle in the haystack, but much more often statistics is about finding the haystack given the needle. One tool statistics uses to get to the haystack is probability.
In statistics there are latin letters and greek letters. When you see a symbol denoted as a greek letter then that is a population parameter. When you see a latin letter that is a sample estimate. It could be Frequentist, Bayesian, Likelihoodist, Fiducial, Empirical Bayes, etc. Theoretical population greeks or sample calculated latins.