Holes in Bayesian Statistics
statmodeling.stat.columbia.edu
statmodeling.stat.columbia.edu
Is the point to end up with an "unassailable subjective distribution"? I believe the power of Bayesian thinking is that you can take a subjective prior, which is necessarily assailable in its subjectivity, and then combine it with data. The result is something that is better than either the subjective prior alone or the likelihood estimator gleaned from data alone.
To borrow a bit from a different paper and anonymous reviewer on one of my papers, inference serves different aims or philosophies. Sometimes it serves more of an estimation function, to increase information about some quantity, or to improve the estimate of that quantity. But sometimes it serves an evaluative, competitive function, in the Popperian sense of affording risky tests of one or more theories or models.
In this latter Popperian aim of inference, priors are to be minimized, which is in many ways the opposite scenario that is assumed with standard subjective Bayesian methods. And even with the former "estimation" inferential aims, there may be situations where you truly have no information or don't feel comfortable assuming it.
What's nice about Bayesian statistics is it still provides a framework for this scenario, in the form of reference priors, in that if nothing else your design and model supporting the parameter(s) to be inferred about implies some kind of assumptions about what you're making inferences about. That in turn can be transformed into a "least informative" prior. So it allows an objective Bayesian framework.
However, in that framework, in many cases you're still often left with uniform priors, which then reduce to frequentist statistics. And in a broader sense, frequentist methods are even further removed from making assumptions in that they completely eliminate the prior from inferential consideration.
There's a tension then, in that in small samples your priors will bias your estimates. If you use least informative priors, you're often doing something akin to frequentist methods anyway. And in large samples the likelihood dominates the posterior so it matters still less.
From a certain perspective, ultimately with Bayesian methods you're making a bet that your priors are accurate enough that the increased bias in estimates will be small enough to be offset by decreased variance due to use of a prior. It's a gamble though, the risks of which will probably vary depending on the costs and benefits of different types of error.
It's nice to see a paper trying to be honest about the problems with Bayesian inference, as it's a bit overhyped at the moment imho.
Several commenters have made this claim, which seems to be true only for maximum a posteriori (MAP) estimates. Frequentist methods do not construct a distribution over parameters.
I disagree strongly - there’s a huge difference in the kind of inference which is provided by each paradigm. To oversimplify, frequentist statistics constructs decision rules with good properties concerning unknown (but fixed) parameters, while Bayesian statistics uses probability to directly reason about those parameters. Confidence and credible intervals are not the same thing - you’re actually flipping what’s considered fixed and random.
(Why even consider MAP when it's dumb? Only because it's computationally cheaper than anything else.)
It’s valuable when you have a combination of useful prior knowledge and insufficient data to review. But, having useful prior knowledge can be difficult to distinguish from useless prior knowledge. Incorrect belief that something is safe means more people are harmed before you update that belief.
This means you can’t say if Bayesian think is a net positive or a net negative abstractly.
This is effectively pure math so there is nothing to disagree or get upset about. What’s up for debate is how accurate your priors are in the real world.
Bayesian stats is nothing more than a rigorous way to transform beliefs + data into a posterior. Yes, flat priors are not always the best choice. That’s not a criticism of Bayesian stats. It’s a statement about how actually formulating a prior is often times the hardest part of a problem. Is Bayesian stats useful for describing or understanding QM? Idk, again, not really it’s job...
Use Bayesian stats, not with an air of suspicion, but a respect for the fact that it will give you the results implied by your data and prior, under your model assumptions. Nothing more, nothing less. By the way, what is the alternative if you find yourself in a situation where your result depends strongly on your prior and you aren’t really sure how to choose your prior? Wave your hands and find an ad hoc frequentist approach? How about just admitting to yourself that your data isn’t enough to make up for the fact that you can’t really quantify your true prior belief?
If you read this and disagree, I sincerely implore you to comment — I just don’t understand the “debate” aspect of Bayesian stats. Some people misunderstand what Bayesian stats is sometimes (very understandable) but I have yet to see a legitimate philosophical or mathematical critique of the Bayesian approach that really made any sense to me. I would like to know if I am wrong though...
I can't tell from your comment if you're aware that Andrew is the first author of the leading textbook on Bayesian methods. http://www.stat.columbia.edu/~gelman/book/
Now that I've taken some time to actually READ the article referenced in this post, I have a couple more things to say; basically: I agree with virtually everything that they're saying in the paper (except that I still don't understand why Bayesian stats failing in QM says absolutely anything about the validity of Bayesian stats), and I think I set up my own straw man here a bit. I thought this was another "Bayesian stats is bad because priors" post, which in a way it SORT of is, but their arguments are a bit more on the side of "you should take 'Bayesian' analyses with a large grain of salt sometimes because...well...priors" <- and THAT I completely agree with.
The only beef I have with these types of articles are when they seem to imply that because priors sometimes have an outsized effect on the results, we should not use Bayesian statistics/analyses. But really, there is no other option. If you have a statistical question, you're stuck with having to answer for your priors. That being said, what IS good advice is to not simply believe every 'Bayesian' conclusion for exactly that reason -- that result is only as good as the (1) prior, (2) model assumptions, and (3) the data. (2) and (3) are easy to fight about, but (1) can be hard to evaluate, and easily can lead all of us astray even when the prior sounds reasonable (flat priors).
Bayesian statatistics doesn't give a satisfying answer/method for choosing priors. Choosing priors is part of statistics. Therefore Bayesian statistics has a "hole" in it.
I absolutely agree with you. Priors are the (very difficult) job of the practitioner.
> choosing priors is a part of statistics
Ah so I think I agree, maybe would nitpick and say priors are part of statistics and CHOOSING priors is part of modeling maybe? Not totally comfortable with that characterization either it’s just that “priors” and “choosing priors” are two different things.
> therefore Bayesian statistics has a “hole” in it.
Completely disagree; this is my main understanding of the argument against Bayesian statistics, I just don’t understand how this follows from the last statement. Priors are hard. But they are explicit! And a necessary ingredient of the posterior. It’s not like we can say “oh priors are hard, so let’s not do Bayesian stats but instead let’s do X” what is X?
I don't know why, but I kept imagining non-local scenarios and thinking, "This seems like it should be OK, so I don't understand what's going on". Having it spelled out is tremendously helpful. I still don't understand what's going on, but at least I don't feel completely crazy ;-)
QM is the number-one favorite topic of crank scientists and pseudoscience. When using it in the discussion of a seemingly-unrelated topic, authors should take extra care to motivate why QM is relevant.
(Of course probability theory and QM are not unrelated topics, but probability exists independently of QM.)
> The second challenge that the uncertainty principle poses for Bayesian statistics is that [...] we routinely treat the act of measurement as a direct application of conditional probability.
Furthermore it states that this problem might also arise for other applications of Bayesian statistics:
> If classical probability theory needs to be generalizedto apply to quantum mechanics, then it makes us wonder if it should be generalized for applicationsin political science, economics, psychometrics, astronomy, and so forth. It’s not clear if there are any practical uses to this idea in statistics, outside of quantum physics. For example, would it make sense to use “two-slit-type” models in psychometrics, to capture the idea that asking one question affects the response to others?
Somehow I think the most fundamentally damning critique, and causality shares this problem, is also the most vague. That applied scientists/experimentalists look at the "automation" that these approaches are supposed to enable and say "that's either doing the trivial part of the job or giving you BS answers".