Stan is a state-of-the-art platform for statistical modeling
mc-stan.org
mc-stan.org
https://www.youtube.com/watch?v=4WVelCswXo4&list=PLDcUM9US4X...
Personally, I think https://probmods.org/ is an exceptionally good introduction to probabilistic programming for someone that knows CS or just some programming and likes a SICP-like textbook that goes into the essence of the topic.
Learning Stan is great, but not as a first probabilistic programming language, because it's quite limited (it trades model expressiveness for performance). So you can't represent a large set of models, such as infinite mixtures, which may become really relevant in the future developments of deep learning. It also has poor performance in models that involve many discrete variables.
For people wanting Python, Jupyter notebooks with Python code examples are here:
* https://github.com/pymc-devs/resources/tree/master/Rethinkin...
Often you hear that deep learning is best at unstructured data (images, sound and recently raw text) and boosted trees / XG boost for tabular data.
Stan users are more often R users than Python user and mostly come from science backgrounds. They often use Stan via a package called BRMS which stands for "Bayesian Regression Models using Stan" which should give you some idea of its core use case.
You wouldn't use Stan if you weren't trying to model your problem as a distribution based probabilistic model.
Deep Learning: Predict an outcome variable
For example, if I want to know what effect household income has on a student's chance of getting into college, Stan would allow you to estimate that given a proposed model.
If instead I wanted to predict a given student's chance of getting into college, I might use Machine Learning.
Of course, those two problems are linked, but it's a fundamental difference of focus.
I'm not saying it is better than DL by any means, as DL can scale much better. Just that I don't think it's necessary to pigeonhole Bayesian inference to just predicting the parameters. In my opinion the "fundamental difference of focus" is just a personal decision, not something inherent to the method.
Can't this be estimated via bootstrapping?
[Edit] I don't understand these downvotes.
The main difference is that with Stan you think in terms of random variables and distributions (and their transformations), while with Tensorflow/DL you think in terms of predicting directly from data. Stan lets model a problem with probabilities and do arbitrary inference, generally asking any question you want about your model.
There are many other interesting alternatives, e.g. http://pyro.ai/ which takes a yet another approach merging DL and probabilistic programming with variational inference. (Stan and TFP can do variational inference too, but I guess it's like Python vs JavaScript vs Ruby vs Java - all of them can be used for programming, but not the same way).
[1] https://pymc-devs.medium.com/the-future-of-pymc3-or-theano-i...
Looks like I get the best of both worlds now.
The other distinction is between discriminative and generative models. In a discriminative model, the output/label is being predicted based on the input features: p(y|x, theta). For example, the probability of an image containing a dog, y based on pixels, x. Theta here refers to the parameters one needs to discover.
In a generative model, one instead models the distribution p(x|y, beta) i.e. given the label, say dog, predicting the joint distribution of all the images.
Neural networks with backproagation can be used for both discriminative and generative models. Bayesian methods can be applied to both discriminative and generative models to compute the full posterior distribution of the parameters, theta and beta.
Edit for clarity: The claim is that the choice of the model vs the choice of inferential methodology (Bayesian vs max likelihood for example) are orthogonal choices.
A neural network doing (discriminative) binary classification based on cross-entropy is maximizing likelihood instead of maximizing the posterior. Most Bayesian examples seem to specify a generative model (a Hidden Markov Model for example) and then infer the posterior. But there's nothing preventing one from using Bayesian methods with discriminative models (generalized linear models) or max likelihood with generative models.
Most particularly, Bayesian modeling tends to be generative modeling as opposed to discriminative. This means that you construct your model by describing a process which generates your observed data from a set of latent/unknown quantities.
For instance, we might observe that n[u, d] clicks are observed on user u on day d for various choices of u and d. We could build a variety of generative stories here: that n[u, d] is independent of u and d, just being a random draw from a Normal(mu, sigma) distribution; that n[u, d] incorporates another unknown parameter p[u], the user's propensity to click, and then is a random draw from Normal(mu + b p[u], sigma); or that we also include season trends sm[d] and ss[d] to both the mean and spread of n[u, d], saying it's Normal(mu + b p[u] + sm[d], sigma * ss[d]).
In these examples, the unknown latents are parameters like mu, sigma, and b as well as any latent data needed to give shape to p[-], sm[-], and ss[-]. Once we've posited the structure of this generative model, we'd like to infer what values those latents might take as informed by the data.
This is the bread and butter of Stan modeling. It lets you describe these generative models as a "forward" process where we sample latents in a simple forward program. Similar to Tensorflow/etc Stan extracts from this forward program a DAG and computes derivatives, but instead of simply maximizing an objective function through backdrop, Stan uses these derivatives to perform a sampling algorithm over the latents (mu, sigma, b).
Ultimately, this gives you a distribution of plausible latent configurations given the data you've observed. This distribution is a key point of Bayesian modeling and can provide a lot of information beyond what the objective-maximizing value would. As a simple example, it's trivial from a Bayesian output distribution to make statements like "we're 95% confident that mu > 0.1".
2) You want to do statistical modelling, not a black box. You already have a statistical model in mind, you just want to fit parameters.
Stan is probabilistic programming system. You describe the data-producing mechanism (the model of reality), and the level and form of approximations used in the estimation. The compiler generates code for the estimators.
However Bayesian inference is a good choice for prediction when you have few data points (deep learning is sample-size hungry). And it is especially good when you have high uncertainty in your labelled training data (ie large variance in the response variable for given input). Here a Bayesian regression (or even classification) model wouldn’t magically remove the uncertainty but rather you’d be able to account for the predictive variance (instead of being none-the-wiser using just good ole deep learning). You can then take it from there how you wish to treat the predictions, given the predictive variance as well.
I don’t see any clear distinction between machine learning and statistics. Machine learning is a type of statistical model which relies on iterative optimization.
Bayesian inference on the other hand is a specific type of statistical model where the aim is to model distributions, not just the output variable (which is what ML is traditionally focused on).
And yes there is overlap, you can take a Bayesian approach to machine learning and that can make total sense sometimes.
You can read more at https://www.fharrell.com/post/stat-ml/ (Frank Harrell is a top statistician who was once a frequentist and now is a bayesian. He writes also on the differences between ML and SM and how to choose between the two)
Others have commented on the role of inference/estimation, and prediction in small data or non-black-box contexts, so I’ll just add that there are deep theoretical reasons to do Bayesian inference. It’s a framework grounded firmly in decision theory, and provides a coherent way to reason about the world. You can prove, under sensible axioms, that beliefs can be described in terms of probability distributions, and that we should update beliefs based on Bayes’ Rule.
But there is a case where it makes sense: when you have reasonably few parameters, and you can understand their meaning. And this is exactly the case of what's called "statistical models". That's why STAN is called a statistical modeling language.
How is that? Gradient descent for these small'ish models is just MLE (maximum likelihood estimation). People have been doing MLE for 100 years, and they understand the ins and outs of MLE. There are some models that are simply unsuited for MLE; their likelihood function is called "singular"; there are places where the likelihood becomes infinite despite the fit being quite poor. One way to fix that is to "regularize" the problem, i.e. to add some artificial penalty that does not allow the reward function to become infinite. But this regularization is often subjective. You never know when the penalty you add is small enough to not alter the final fit. Another way is to do Bayesian inference . It's very slow, but you don't get pulled towards the singular parameters.
I've personally liked PyMC for simple models and relative ease of inference, as it's more integrated with the Python language. That being said, if you want the latest in inference methods and statistical alchemy, Stan is the place to go.
You do still need the C++ toolchain, but can just write your code in R.
A Stan program can be run in any of our interfaces in Python, Julia, R, MATLAB, Stata, etc. But you can't mix any of those languages into a Stan program.
The C++ toolchain is required because Stan transpiles its programs to C++, then compiles those against the Stan math librarym, which does autodiff. But you don't need to write any C++ to use Stan, just to develop extensions for it.
Were you?
Bravo to the team behind it and for making and supporting such a powerful tool!
Here is an example from a class taught last January that uses stan to fit a simple ODE (using the `integrate_ode_rk45` function in stan):
https://github.com/gregbritten/BayesianEcosystems_IAP/blob/m...
likewise as far as I know, linear regression and mixture models are both done in a Bayesian style (a hierarchical model giving priors for parameters).
As currymj says, the differential equations (same for all the linear algebra solvers like eigendecomposition) can be used in defining likelihoods for either Bayesian or frequentist estimation. Same for all of our linear algebra operations and special functions.
Not every model that can be programmed in Stan has a well-defined MLE or proper posterior. Standard hierarchical/multilevel models don't have MLEs, even with standard shrinkage. Bayesian models with improper priors and no data wind up with improper posteriors, etc.
Having said all that, almost all of the use of Stan is for Bayesian inference.
https://benchmarks.sciml.ai/html/ParameterEstimation/DiffEqB...
https://benchmarks.sciml.ai/html/ParameterEstimation/DiffEqB...
It also likely depends on how well conditioned your model is - even if you can get it to run for huge models on reasonable hardware, convergence may not be practical.
Stan can parallelize multiple chains and it can parallelize the density/gradient calculations in a single chain. But for the latter to be efficient, the chunks being parallelized need to be compute intensive, like you might get in a pharmacometric compartment model where you might have to solve a bunch of differential equations for each of thousands of patients in a clinical trial.
I find there's a lot of excitement around Bayesian inference and MCMC, but I do wonder about the substance.
Stan fittings can be made parallel, some models will scale linearly with the data, but in the main you won't find many big data use cases here.
You also can't use Stan for online learning.
Can’t you loop posteriors as next iterations priors to get a system that learns online?
In the case that your model is extremely simple such that your posterior has a "conjugate prior" (i.e. the posterior and prior are the same family of distribution), this sort of loop back is possible. But where this is possible you have no reason at all to use Stan or MCMC since you can just update your posterior directly.
Sampling is a slow approach when there are other alternatives. For example, if you are after OLS regression, you can do the equivalent with Stan but it may be an order of magnitude slower than plain OLS. Further, the calculation of your likelihood function will scale linearly with the size of the data. But adding new parameters will scale exponentially, so you may find that a model with 2 free parameters which takes 10 minutes to fit takes 2 hours with 3 parameters.
A good thing about Stan however, is that it is parallelisable so you can run it on many cores (and it will scale linearly for a good while) and you can also run it on MPI across many machines. Some regression functions with very large matrices support GPUs (although Stan requires double precision to work). So to some extent you can "throw more money at it" to get a result out and it has been used for very big data problems in astronomy for example which however utilised something like 600k cores if memory serves correctly.
Adding new parameters scales as O(N^5/4) in HMC, whereas it scales as O(N^2) in Metropolis or Gibbs. It's quadrature that scales exponentially in dimension. There's also a constant factor for posterior correlation, which can get nasty. I regularly fit regressions for epidemiology or genomics or education with 10s or even 100s of thousands of parameters on my notebook with one core and no GPU.
MCMC or optimization can be sub-linear or super-linear in the data, depending on the statistical properties of the posterior. Some non-parametric models like Gaussian processes can be cubic in the data size, whereas regressions are often sub-linear (doubling the data doesn't double computation time) because posteriors are better behaved (more normal in the Gaussian sense) when there's more data and hence easier to explore in fewer log density and gradient evaluations.
If by "big data", we're talking about too big to fit in memory, that's right. Stan's fully in-memory. Compute can be distributed and GPU-powered for matrix ops, but all of the data and parameters and the core autodiff expression graph need to fit in memory.
For "medium data", Stan's adaptive Hamiltonian Monte Carlo sampling is much more efficient and scalable to complex models and higher dimensions than Gibbs or Metropolis. I'm fitting a Covid prevalence model using a custom trend-following and mean-reverting second-order autoregression model over 400 distinct regions with weekly data that has 5M data points and 10K parameters and adjusts for sensitivity and specificity of various tests taken. It fits in a single thread using MCMC in 24 hours or so, but we can fit the model with variational inference in a couple minutes. Although variational inference often produces reasonable point estimates in bigger data settings, it doesn't reasonably quantify uncertainty. I'm also working on a genomics model for differential expression of splice variants that involves 120K measurements and just as many parameters to deal with overdispersion of biological replicates in a control and treatment group. We're using variational inference and it fits in a couple minutes for the comparitiver event probabilities we need to estimate.
These methods are ideal for small datasets with correlation structures that aren’t necessarily independent.
Also great if you want uncertainty with your estimates.
There are hundreds of different applications of Stan across the physical, biologial, and social sciences, as well as in finance, education, sports analytics, actuarial sciences, transportation planning, all sorts of material and chemical and civil engineering, clinical trials and pharmacometrics, etc. etc. It's most popular in fields like ecology and epidemiology where Bayesian methods are already popular. For instance, many of the Covid models (like the one for NY state) are being built with Stan. All four baseball teams in the semifinals (LCS) use Stan for analytics, for example. Google and Facebook use Stan for ad attribution and resource allocation. It's been used for models of neutrino mass and models of galactic mass, models of supernovas, and it's even used in the LIGO gravitational wave experiments.