Bayesian statistics and machine learning: How do they differ?
statmodeling.stat.columbia.edu
statmodeling.stat.columbia.edu
Generative inference: We know the process that generates our observed data and can model it to sufficient precision, s.t. we can establish a forward synthesis: Parameters in, synthesised data out. With that we can compare a synthesized dataset from a configuration of parameters with real observed data and, using Bayes (a.k.a. MAP estimation), infer the correct configuration of parameters that most likely generated our observed data, which is the answer we're looking for. Examples: Kalman Filters, Linear Regression.
Discriminative inference: We do not know the process that generates observed data or it's too difficult to model. In that case we ask machines to take a shot at modelling it. To do that we set up parametric data transformation pipelines (e.g. Neural Nets) and feed it lots of correct input/output pairs. Over time the model will learn how to transform the input (observed data) to the output (answer we're looking for) and hopefully generalise well when we feed it new input data that it hasn't seen before. Examples: NNs for classification.
Of course there are interesting mixtures and hybrids between the two and the lines are blurred in some cases but this is the general distinction.
With ML, the utility of the model is largely a function of its predictive power. You want to accomplish some task. "Make it go"
With much of statistics in a research context (such as the social sciences as called out in the link), the interest is more on explanatory power of the independent variables. Most social scientists would happily trade a bit of predictive power for a more explanatory model that neatly maps back to a set of hypotheses. "Why does it go?"
[0] https://www.stat.berkeley.edu/~aldous/157/Papers/shmueli.pdf
Machine learning fundamentally cares about model performance on other data, like validation or test sets. It is looking for models that perform well, but not necessarily model the underlying process.
Bayesian statistics, like most statistics, wants to accurately estimate parameters in the model. It cares most about models that are portraying the underlying data-generating process.
When a model for a binary outcomes returns 0.9 for a given data point, that implies a 90% probably that the value is true.
Evaluating the quality of this estimates (often called measuring the calibration of the model) is even very common.
(There are some exception of course. Max margin models aren't probabilistic. And sometimes people use fixed variance parameters for their normal models, etc).
[0] https://www.sciencedirect.com/science/article/pii/S156625352...
See my other comment for more detail.
You can actually recover from this a bit. I saw a paper once where they used the Hessian to approximate the posterior as a gaussian distribution around the maximum likelihood. Can't remember what this paper was called unfortunately.
Consider tossing a coin. If I see 2 heads and 2 tails, I might report "the probability of heads is 50%". If you see 2000 heads and 2000 tails you'd also report the SAME probability estimate -- but you'd be more certain than me.
Neural networks give probability estimates. Bayesian methods (and also frequentist methods) give us probability estimates AND uncertainty.
The literature on neural network calibration seems to me to have missed this distinction.
[0] https://jeffreyling.github.io/2018/01/09/vaes-are-bayesian.h...
I'm just trying to point out that there is a dichotomy between the Bayesian and the non-Bayesian, and that the standard neural network models are non-Bayesian, and that we need Bayesianism (or something like it) to talk about (epistemic) uncertainty.
Standard neural networks are non-Bayesian, because they do not treat the neural network parameters as random variables. This includes most of the examples that have been mentioned in this thread: classifiers (which output a probability distribution over labels), networks that estimate mean and variance, and VAEs (which use Bayes's rule for the latent variable but not for the model parameters). These networks all deal with probability distributions, but that's not enough for us to call them Bayesian.
Bayesian neural networks are easy, in principle -- if we treat the edge weights of a neural network as having a distribution, then the entire neural network is Bayesian. And as you say these can be approximated, e.g. by using dropout at inference time [0], or by careful use of ensemble methods [1].
[0] https://arxiv.org/abs/1506.02142
Quote: "Deep learning tools have gained tremendous attention in applied machine learning. However such tools for regression and classification do not capture model uncertainty."
[1] https://arxiv.org/abs/1810.05546
Quote: "Ensembling NNs provides an easily implementable, scalable method for uncertainty quantification, however, it has been criticised for not being Bayesian."
Bayesians use the terms 'aleatoric' and 'epistemic' uncertainty. Aleatoric uncertainty is the part of uncertainty that says "I don't know the outcome, and I wouldn't know it even if I knew the exact model parameters", and epistemic uncertainty says "I don't even know the model".
Your example (outputting a mean and variance) is reporting a probability distribution, and it captures aleatoric uncertainty. When Bayesians talk about uncertainty or confidence, they're referring to model uncertainty -- how confident are you about the mean and the variance that you're reporting?
I’d hazard a guess that analytical solutions are intractable and numerical solutions would be infeasible.
I personally believe "machine learning" is a term that could be removed from use and nothing would suffer (and clarity might even be improved). I feel somewhat similarly about "data science" although at least that captures some intersection of database engineering and computational statistics that is useful when discussing very large data problems.
yes
> would have said "machine learning" meant "deep learning models"
no not at all.. DeepLearning has gotten serious technical press since about 2015 when certain competitions yielded better-than-average-human performance. DL is also closely coupled to the business model of massive cloud providers who handle streams of digital media. Yet ML stats continue to reliably predict eighty-percent plus well, on lots of kinds of problems.
> it seemed to mean "computational multivariate discrimination and classification"
no, classification has always been on of two main uses for ML, the other being regression-style prediction of real values.
> believe "machine learning" is a term that could be removed from use
no, stats are useful and continue to be useful. It is AI that is the really wonky term, to my ear.
Please consider that there is a lot of difference between academic statistical methods, and technical press product hype. You may be suggesting a change in media coverage jargon, but you know, good luck with that..
The early ML models used SVMs, Logistic Regression and Naive Bayes. Not deep at all. What made them ML models was the size of the feature sets data, and the use of automated feature selection.
If you are curious, the book "Probabilistic Machine Learning: An Introduction" by Kevin Murphy goes into greater detail.
For example in Reinforcement Learning P(R | A, S), eg. probability of reward given action and state, while in Bayesian network P(A | B, C), probability of even A given B, C distributions, in such networks you can do inferences.
I have no idea how correct or false this explanation is.
"Before we started calling it nanoparticles, we used to just call it chemistry, but then we realised we got a lot more citations that way."
[0] https://scikit-learn.org/stable/modules/mixture.html
[1] https://scikit-learn.org/stable/modules/gaussian_process.htm...
> Written in a clear and concise style, Modeling Mindsets introduces approaches such as Bayesian inference, supervised learning, causal inference, and more.
> After reading this book, you will have a much better understanding of the different approaches to modeling and be able to choose the right one for your problem.
https://book.modeling-mindsets.com/
Edit: Ah darn, I forgot about this part: "You should feel comfortable with at least one of the mindsets in this book". So perhaps not the best start if you don't have a base in at least one method. For frequentist statistics, consider https://www.openintro.org/book/os/