May Bayes Theorem Be with You
technology.stitchfix.com
technology.stitchfix.com
Bayesian statistics abandons this worst-case approach, instead opting for an average case analysis. Here, we average over the all the possibilities, weighting their relative merit by the prior. The analysis is always a little bit conservative (thus the connection to regularization), but it is never "worst case" in the same way that classical statistics operates under.
Lots of the other talk in this article is not really about classic vs Bayesian statistics at all. Both methodologies are perfectly happy working with more complicated, hierarchal models. Both approaches have plenty of work dealing with regularization, and both methods will suffer if you mis-specify your model. The fact that Bayesian analysis is less likely to "crash" in a case of model mis-specification can be thought of as just as much of a drawback as it is a benefit.
It's always seemed to me that the differences I see, such as in ML vs. MAP, are not due to a question of worst-case versus average-case analyses, but rather no-prior-knowledge versus yes-prior-knowledge analyses. I've never seen a proof that classical statistics gives a worst-case bound.
The worst-case nature of classical/frequentist statistics is very well exposed in PAC (probably approximately correct) analysis:
http://en.wikipedia.org/wiki/Probably_approximately_correct_...
In real life, both the parameter value and the data are fixed. If you knew the heights all of the humans on earth, you could compute the exact mean human height (at the time). This mean is the true mean, it's the parameter value we're looking for, and it's absolutely fixed.
Maybe one is more intuitive than the other, but I'm unswayed by the "one is fixed in real life" argument.
I see your point. Thanks for the input.
I do think that P(hypothesis | Data) - the Bayesian way - is more intuitive than P(Data | hypothesis). I also think that the confidence interval is often misinterpreted as a credible interval. But that does not mean that the credible interval is the right choice for every analysis.
From a purely practical sense, the data is fixed when you have your sample or have collected your time series data. This is the data you have and you can observe it. I do agree that the data may change when you collect more samples or get more history. But the Bayesian framework is set up well to handle that, as today's result can be tomorrow's prior.
Having said all this, I don't think this debate takes away from the use cases outlined in this post. This is not meant to read as "Bayesian versus frequentist" but rather to highlight some important benefits of Bayesian analysis. For example, if I am building a pricing model that will be used to make actual decisions and my classical model is spewing out strange coefficients that do not seem to be consistent with external data and common sense. I want to infuse that model with outside information before I use the model to make real life decisions, or just shrink it. Bayesian regression gives up a structured and transparent way to handle these types of situations.
"The data is fixed" refers only to the data we know. (Otherwise, as you point out, there would be no issue since we would simply calculate the population statistics directly.) In the frequentist model, we have to pretend that this fixed data is actually a random distribution in order to calculate the probability we're interested in. In the Bayesian model, we just combine the known data with the prior to get the posterior probability; we don't have to pretend anything.
Okay, yeah. The sample is fixed. The entire population is random. I'll use this terminology.
> In the frequentist model, we have to pretend that this fixed data is actually a random distribution
With confidence intervals, we're modelling the entire population as a random distribution, using the fixed sample's mean and variance to compute estimates of the population's mean and variance. I think this is different than "pretending the sample is random".
If you constructed a normal model of the sample, you would just use the sample's mean and standard deviation. But to model the population, you typically use the sample's standard error as an estimate of the population's standard deviation. This is critical. You have to account for the fact that the population is much larger than your sample.
My knowledge is pretty limited, I have heard of BEST but happily t-test away on a daily basis. Essentially what I'm looking for is a "I know t-tests and ANOVA and use them regularly...how would I switch all that to a Bayesian approach".
Does the book I look for even exist or would my best bet be reading the BEST paper (and the author's website/youtube video)? Edit: It lookes like the second edition of "Doing Bayesian Data Analysis" by Kruschke looks good. It does have a dedicated chapter on NHT (vs. MCMC)
http://nbviewer.ipython.org/github/CamDavidsonPilon/Probabil...
Does anyone know if it's an available theme, or if it was custom built?
For problems where you have small amounts of data but a lot of intuition about how the where the data comes from, then Bayesian methods will typically give you a better answer. However, if you have a small amount of data and a bad prior, then your conclusions will be flawed. Essentially, it's garbage in, garbage out.
Naive Bayes, for example, is more of a "machine learning" technique where the goal is to classify people into groups based on features. Naive Bayes is called Naive because it assumes that all regressors (x_j) are independent given the target variable (let's call it y and assume it is binary). In other words, the conditional log odds of y=1 given the x_j variables is equal to the sum of the log density ratios, where the log density ratio for variable x_j is ln(f(x_j|y=1)/f(x_j|y=0)).
On the other hand, in the price elasticity example described in post we want to infuse outside knowledge into the model because we don't believe what it says on its own. This is a situation where interpretation and believability is an important part of the objective function because we will be running future pricing scenarios from the model.
If you are building, say, a churn model to predict who is going to cancel their accounts, you probably wouldn't infuse your model with outside knowledge since cross validation accuracy is your main goal. You might regularize your model, however, which can be done in a number of ways (Bayesian or non-Bayesian). But in a pricing model or media mix model, and many other cases, the use case above is very real.
I suggest reading the “Elements of Statistical Learning” by Hastie, Tibshirani, et al.
I'm just picking up Machine Learning these days and this is a good Bayesian intro.