Bayesians moving from defense to offense
statmodeling.stat.columbia.edu
statmodeling.stat.columbia.edu
- When sample size grows, frequentist and bayesian (if the prior is not too restrictive) point estimates seem to converge to each other anyway
- The distribution of your point estimate (frequentist) vs. the estimated distribution (bayesian) also don't seem to differ too much either
- When the sample size is small the Bayesian prior dominates
- Interestingly, when I see Bayesians simulate random data (to introduce the concepts on this data) they usually assume a true parameter value. E.g. when sampling from Y = a + b * X + e, they'll assume fixed, true values of a and b and not random variables - which is a frequentist assumption! So far I've never seen e.g. b being sampled from Normal(mu=2, sigma=1) instead of just setting b=2.
- The frequentist assumption of a true population value which we try to estimate just makes sense to me. For example there is a true mean income over the working population. It's not a random variable but a fixed value which can be computed if we just asked every single working person for their income and then compute the mean over all values.
I tried getting into Bayesian stats but honestly it just seems overkill for most cases. For a simple regression computing b_hat = inv(XX')Y is just faster and easier than numerically sampling traces. Bayesian forces you to think about the data generating process - I appreciate that, but you need to the same when it comes to frequentist stats, it's just a little less obvious.
Yes. And so? Bayesians would argue (and I quote) that "the interesting limit in statistics is when the number of samples tends to one. The limit when the number of samples tends to infinity is completely useless."
> I tried getting into Bayesian stats but honestly it just seems overkill for most cases.
There are 3 black balls and 7 white balls in an opaque bag. How likely is it to pick a black ball? Bayesian statistics gives a straightforward answer (you just assume an uninformative prior and perform a computation). But frequentist statistics starts to argue about an infinite number of replicas of your own universe and other nonsensical constructions. Not sure that the Bayesian approach is overkill in that case...
The "and so?" is answered right after that. The prior dominates, which is a bad thing.
The smallest amount of samples you can use is 1, isn't it? If you have 0 samples then you do nothing because you have no data. Is there a way to have half a sample?
> if course your belief tends to whatever your belief was before you saw any data
Your beliefs should tend to that, sure, but if you're trying to produce an actual number for sharing then your beliefs shouldn't be a huge factor, and an uninformative prior being a huge factor is also bad.
For numbers that leave my head/notebook, I'd rather keep the new evidence by itself and say it's weak.
I'm not sure if by "absence of belief" you mean "ignorance" or something else.
If you have a die and you don't know anything else about it you should assume that the probability for each side is 1/6.
If you also know that the expected value is 4 (instead of 3.5 for a fair die) there is a way to calculate the probability distribution that reflects that constraint - and nothing else.
Now, if you don't even want to think about anything Bayesians can do that too.
On the negative side, the frequentist approach doesn’t produce a post-data probability for the thing of interest either.
It provides the probability of something else - as you mention - which can also be interesting but it’s not what people really would like to know (as the generalized misinterpretation of the meaning of frequentist results makes clear).
It would be weirder if the result didn’t depend on the things assumed.
I don’t know what kind of questions are you thinking of but outside of mathematics they are rarely fully specified.
If the answer changes enough depending on the additional assumptions to seem weird that is a sign that the question was not completely clear.
Of course Bayesians can also say that there is not enough information to provide an answer when that’s the case, just like they can make additional assumptions explicit to provide one.
I know it's not like that. But it's still weird that at the end of Bayesian analysis the best you can deliver is if-by-whiskey style deliberation.
I know it's still valuable. Just weird.
The thing with non-Bayesian analysis is that they don’t answer at all the question “what’s the probability of X conditional on the data observed”.
Let's say you have a personal belief that something is going to happen with probability x. Would you actually want to tell others that the probability is y, because that's what the data says, without letting people know that for other reasons that are not reflected in the data, you truly believe it is x?
Your informed opinion incorporates this dataset, but you shouldn't imply it's "based on" this dataset.
But you shouldn't share a frequentist parameter estimate or confidence interval if you have prior information that would influence it non-negligibly, at least not without sharing that prior information also.
Yes frequentist statistics work very well in practice, but it's a bit adhoc and suffers from various problems like say if you estimate velocity and estimate kinetic energy, you get values that are incompatible which is kinda ugly and non-intuitive and makes you want to dig deeper into how such a thing happened.
Bayesianism has the answers.
Also sometimes it really does matter like in medicine, where some conditions have a very low prior probability.
That's how many people feel about Bayesian methods when trying to pick an initial prior.
The Bayesian philosophy of "random parameters" does not mean that Bayesian methods cannot be assessed for frequentist properties or compared against frequentist procedures.
Multilevel models are fantastic to address a problem that is often ignored by frequentist approaches, the need for shrinkage and information sharing. This pops up all the time in modern statistics. For example, if you test 1000 hypotheses, calculating p-values and adjusting these with some multiplicity correction scheme is not sufficient.
You should borrow information across random variables with a multilevel model to avoid estimating exaggerated effects in tests whose outcome is deemed to be significant. Andrew Gelman's post is concerned with this topic.
Another point is that Gelman et al. use weakly informative hyperpriors. These are not really subjective. If anything, they usually regularize solutions by pushing effects towards zero. Plus, on multilevel models, priors are only needed on hyperparameters.
However it seems that you are suggesting another use. If I have 10 cognitive measures each measured once in my subjectd, the default has been to do a multiple comparison correction, either FDR or FWER on 10 tests. We know that the 10 tests are not truly independent, so Bonferroni is probably too conservative.
It seems here you suggest running this with test being a random effect. I've seen this approach with item level data in a task, but I didn't really think to do it when the tests are not from the same battery, construct. And more to the point, this fixed effect model would be of no particular interest, while random effect CIs are difficult to estimate. So I am left a bit confused.
Ideally one should use the whole posterior distribution of your model parameters which is not the case for point estimates.
>So far I've never seen
Because people are lazy.
Bayesian works great if you have great knowledge in your field and you can fine tune everything. Frequentist stats just works and easily interpretable but easy to make mistakes esp. when starting out.
This is a historical issue because of some hard-headed frequentist founders, but in modern days the frequentist concept of confidence distribution is gaining acceptance, which is the proper frequentist equivalent of the posterior, so this distinction between Bayesian and Frequentist is disappearing.
Rather than giving specific point estimates or interval estimates, calculating a frequentist confidence distribution allows you to compute confidence intervals for all possible confidence levels, just like the posterior does. See this excellent review paper for more info on this: https://statweb.rutgers.edu/mxie/RCPapers/insr.12000.pdf
The key insights is that a confidence distribution is an estimator for the parameter of interest, instead of an inherent distribution of the parameter.
The major distinction remains: Frequentist confidence intervals are something quite different from Bayesian credible intervals. I don't think that having a distribution that can be used to calculate any desired confidence interval - like the posterior distribution can be used to calculate different credible intervals - changes much.
http://www.stat.columbia.edu/~gelman/research/published/pval...
Not everyone wants to admit - or even understands - that the % confidence is about how frequently such intervals will contain the true value when the procedure is applied to other populations / datasets.
"[frequentist confidence intervals] don't really tell us anything about the probability that the statistic of interest for our population is contained in the interval"
and
"the % confidence is about how frequently such intervals will contain the true value when the procedure is applied to other populations / datasets"
My understanding is that a 95% confidence interval implies that there is a 95% chance the true statistic lies within the interval. What are you saying it means?
To some degree I'm interested in the philosophy behind probability but there are limits to my concern with it. If you're saying that ontologically this is not truly a 95% confidence interval, but practically speaking, it does mean that, then I'm more interested in the practical interpretation because it's more relevant to applications of statistics.
If you calculate one thousand confidence intervals in 950 the true statistic will lie within the interval. I'm not sure if that means that there is a 95% chance the true statistic lies within a given interval.
Let me use a extreme example to highlight the issue. Imagine that I use the following procedure to produce a 95% confidence interval for whatever quantity you're interested in:
with 95% probability return the interval [ -1e999 1e999 ]
with 5% probability return the interval [ -1e999 -1e999+1 ]
If you calculate one thousand confidence intervals in 950 the true statistic will lie within the interval. I wouldn't say in any given case that there is a 95% chance the true statistic lies within that particular interval. (There was a 95% chance though.)
This article (open access PDF available) discusses how different methods of producing confidence intervals perform in a toy problem:
The fallacy of placing confidence in confidence intervals https://link.springer.com/article/10.3758/s13423-015-0947-8
One could argue that those misleading flaws can be avoided by being careful with what confidence intervals do you use and how - but the point is that coverage probabilities are about a set of counterfactual situations and not about the actual situation at hand.
https://www.redjournal.org/article/S0360-3016(21)03256-9/ful...
I didn't understand much from the part where they use Markov chains to calculate probability of something that is hard to know but the rest illustrates differences really well.
It is ultimately the same math, if you are comparing apples to apples. Just different ways of looking at a problem. Sometimes one is better suited than the other.
It’s like as if physicists were arguing over whether Cartesian or Polar coordinates were better. It’s the same damn physics, just expressed differently. In some problems one approach is easier to work with than the other, and can even make seemingly intractable problems solvable. But that doesn’t mean the other approach was “wrong.”
The controversial part is the methodology for statistical analysis built on top of it (e.g., you need a prior but where did that come from?)
I will do it:
1. Computationally easier
2. often analytical theory available for most use cases so interpretability is high
3. more literature available so you can get unstuck faster if you mess up
4. no accusations of subjective bias in your prior (the con is clear, no ability to leverage subjective expertise)
5. In the asymptotic regime, MLE and bayesian MAP often converge anyways
6. king of hypothesis testing
For most people, it doesn't matter. It matters when you are doing treatment for small sample sizes or other situations that would cause low power.
This advantage can't be understated. Researchers, at least in psychology, almost never seem to care about the magnitude of an effect. It's simply enough to show that some effect happens in some direction. For this (generally acceptable) purpose, frequentist stats are great.
It's not entirely unreasonable, as often extra variables suppress the effect size. E.g. the effect is substantial in people with some genetics, insignificant otherwise. The fact that there is an effect at all makes it interesting as a starting point for further exploration e.g. to determine the mechanism or find the extra factors needed to make the effect significant.
That's not an advantage, that's one of the primary reasons psychology is among the worst subjects hit by the replication crisis.
So not a cheap shot at all IMO.
For example, suppose you are researching the effect of an emotional/negative picture on how a participant makes decision in an economic game. Here, you may think “a big effect size implies this study will have a large practical relevance.” However, following typical statistical designs, an effect size may be largely determined by just how many trials the experiment used for each participant. This is also not the only factor that’s relevant. For instance, it’s debatable how the emotion induced in the task compares to the intensity of real life emotional situations. With these factors in mind, the effect size really doesn’t tell you much that can be applied outside the laboratory.
Maybe in some non-experimental sub-fields (e.g. personality psychology) effect sizes are more meaningful…
Bayesian statistics is sound but I suspect it's often just used to justify biases. It is technically valid to use a prior and de facto never update it, because you know I'll get around to updating my prior next week, or... eventually, cough cough let's be honest, never
It's the frequentist one that gets your biases implicitly, on the form of corrections and hypothesis formulation, so that people don't notice them.
One issue is that Bayes estimates are almost always produced, even if no information is coming from empirical data, and all the information is coming from a prior. So it's possible to produce results heavily influenced by the prior with Bayesian estimation that with frequentist methods would fail completely because of lack of identification of the model, sending a strong signal that something is wrong. This can all be sussed out with Bayesian methods but people often don't do it.
Another more subtle issue is people aren't quite aware of how a prior can deviate from "maximal conservativism". Sometimes, for example, depending on the model, a very flat prior is actually not conservative, and is overweighting tails.
There's other examples too. Basically, yes, Bayesianism forces you to be explicit with your biases, but people are really bad at interpreting the actual impact of those biases in a formal Bayesian framework, or at least, aren't any better at it than with frequentist methods that are available.
If you approach statistical inference from the perspective of accuracy (as the linked paper seems to do) Bayesianism is better to the extent the priors are accurate. This is true a lot of the time empirically, but it does lead to a kind of tautology, in that you're doing the analysis because you don't really know what "truth" is. So if you're right in your priors, Bayesianism is accurate, but then you didn't really need new data as much in the first place; if you're wrong, it's more biased. Basically in the bias-variance tradeoff, Bayesianism makes a bet on reduced variance assuming that the resulting bias will be small enough.
Philosophically, though, there's a completely different argument, which is one of competitive fairness. You might say this doesn't matter, but consider consequential decisions, like hiring or admissions decisions: if someone was making a prediction about you, would you want them to use a strong prior, or something that's maximally conservative and fair?
This philosophy leads to frequentism basically.
My preference is to be maximally conservative in a Bayesian framework, which leads to reference priors, which are often flat in many canonical situations, which is basically frequentism. In other situations you might have a different kind of prior.
To me the linked paper is pretty interesting and makes a good point. On the other hand, I'd rather not make any assumptions about a new result based on past studies on other effects. I'd rather just collect lots of diverse real data and meta-analyze it. There's no substitute for data -- and that includes priors.
I broadly agree with you, but I'm wondering if you would reconsider your qualification as "less complicated" if you consider beginner learners. E.g. someone who knows basic descriptive statistics and probability theory, and is making first contact with inferential statistics. Specifically, assume a learner who knows what an integral is, but is far from proficient with it (UGRAD student, not a GRAD student).
I was reading this paper[1] recently, which highlights two difficulties of teaching Bayesian stats: 1) the mathematical complexity of understanding conditional probability distributions, and 2) the lack of well defined, broadly accepted conventions for what priors to use in specific data analysis scenarios.
I think a computational approach to prob theory could mitigate 1), but 2) remains a problem—the freedom to choose priors, is also a burden...
[1] https://www.stat.purdue.edu/~dsmoore/articles/BayesPedagogy....
Maybe someone here might have suggestions?
The closest thing that comes to mind is "Bayes factors," which has some traction (usage), but apparently they have lots of problems and limitations too, cf. https://www.youtube.com/watch?v=MqeWpR6S4XA
The canonical approach is to build a generative model with a parameter (or multiple for ~anova) that codes for the difference between groups and do inference on that parameter of interest. Most of the recipes taught in statistics classes can be modelled as a regression of some kind (this counts for frequentist stats too, see https://lindeloev.github.io/tests-as-linear/ ). Some advocate to do that inference with bayes factors. Others, like discussed elsewhere in this thread, advocate combining the resulting posterior with a cost/value function, but either way the lesson is that there is less focus on "t-test-vs-anova" because they're the same thing anyways.
I had previously started the BDA course, which is another famous Bayesian course, see https://avehtari.github.io/BDA_course_Aalto/ but I didn't finish it due to travel.
No more excuses in 2024... time to level-up the Bayesian modelling skill ;)
Of course the winners don't rant about anything. But any time we probe the consequences of Frequentist statistics they turn out to be horrific for science, our health, and our planet.
It’s like how you wouldn’t hear people in the US promote imperial measurements for home baking. When you dominate, you don’t have to advocate.
My college offered stats in two tracks: There was a one-semester course for science majors, which was mostly plugging numbers into formulas. The course was utterly baffling for most of the students who took it.
There was also a two-semester course for math majors. It was mainly about proofs, but also had time to go into more depth. You have to know the assumptions underlying a formula, if you're expected to prove it. ;-) But it was only taken by the math majors.
One thing we didn't have when I took stats was computers. I graduated from high school in 1982, and the colleges in my state were just beginning to get computers. I wonder if stats could be approached differently if it could start with nothing but data -- lots of it -- and graphing tools. This could even happen pre-college.
First course was almost purely frequentist
Second was half and half with several weeks approaching the frequentist bayesian approach from both sides
Our third stats course was mostly bayesian though self directed as to which point of view was “most appropriate” for the formulated problem sets. I remember roughly half the cohort used a frequentist perspective on a fairly obvious (but not explicit) bayesian assignment and were severely punished by the course coordinator (fairly for a late year subject tbh)
You can squish it into Bayes by considering it a uniform prior on the real number line that you never update, but you’re not really doing things “in the spirit” of Bayes, then.
IMO, null hypothesis testing is way overused. We should be quantifying the comparison of the null versus alternative hypotheses. People already compare them. Might as well bring some math into it.
If your hypothesis is that A is greater than B, then you're test boils down to "The probability that A is greater than B", which is arrived at through parameter estimation via Bayes' Theorem.
You can consider the distribution of estimates for a parameter to represent a space of possible hypothesis for the true value of that parameter and their absolutely or relative likelihood based on the information available.
The also has the pleasant consequence that hypothesis testing using Bayesian methods nearly always is closer to the actual question someone wants to answer rather than replying with confusing statements such as "the probability of rejecting the null hypothesis", which contains an implicit double negative.
but even in this situation you would need a threshold for decision making, i.e. a p-value.
Bob bob bob'bob bob bob bob bob, bob bob bob bob bob "Bob".
Bayesian decision making typically involves combining a cost function with your posterior distribution.
It's worth pointing out that this is typically not possible in most frequentist frameworks since they work in double negative style assertions and arbitrary threshold rather than modeling the problem itself.
While you can map Bayesian approaches to frequentists tests, this is usually a misunderstanding of Bayesian methods coming from a purely frequentist background. Bayesian analysis is fundamentally more flexible since it just the application of the sum and product rules (you don't even need Bayes' theorem since it can be trivially derived from these) as well as corresponding cost/value functions.
In the most common situations this typically means the expected value well exceeds the expected cost and you're comfortable with the risk if you're wrong.
A prior could be any previous knowledge that you have on the subject, or anything about the world.
Let me give you a silly example. Imagine I do an experiment, on if person A can see the future, and predict coin flips.
Now, imagine that I do this experiment, and lo and behold it passes the 5% threshold P value!
Remember what a P value is. It is simply the percent chance that the results that you got were by chance.
So, if I do that experiment, and it actually gets at the "P value" of 5%, which means that there is a 5% chance of it happening by chance, do you know believe that Person A can see the future?
I wouldn't. I would need a much lower P value to start believing that.
Thats all a prior is. It is using the fact that "Person A can see the future!" is an extraordinary claim. And for that, we require extraordinary evidence. Much better evidence than a mere "5% that it happened by chance". P
P-values are fundamentally Frequentist because this framework see that observed data as random, whereas Bayesian statistics believe our observations are the only thing that is actually known, and all the other parameters are what is random in our experiments. That is: the data is the only part of an experiment that is real, everything else that we can't directly observe is something we must hypothesize about.
At least in my undergrad education, I took two dedicated statistics classes and many more domain science classes that used statistics, all of which were so steeped in the frequentist paradigm that frequentist vs Bayesian wasn't even mentioned. Frequentist statistics simply was statistics for me, until grad school.
I would love to embrace (or at least try to) such a new approach, but it feels like without a PhD in stats it's hard to get it started.
Some very basic Bayesian models can go a long way towards making informed decisions for a/b tests
There are some caveats though - and I mention these from the experience of running such solutions on a large scale in production. First, BB-MAB can't adapt to context by design. They only look at click/no-click behavior across the population. So, if your population has two distinct segments - youth and elderly - who behave very differently wrt purchases, the BB-MAB won't pick a different winning advt. per group; its blind to these groups.
The solution is to use something like a contextual MAB - which assimilates user features (or whatever you might throw at it) into the MAB. There are simple ways to adapt simple MABs to the contextual setup [2] (in my experience, these can also be effective) but, of course, the literature in this area is wide and deep.
A second caveat is that if the ratio of the size of the pool of advts. to the number of impressions is high, the BB-MAB won't converge or converge to a good optima; the search space is simply too large relative to the data. In cases like this it becomes important to begin with the right Beta priors, instead of the standard recipe of starting with a Beta that looks like a uniform distribution.
But I wanted to say that when I looked A/B testing few years ago I started off with this book.
https://github.com/CamDavidsonPilon/Probabilistic-Programmin...
And somehere down the line he will introduce the Beta/Binomial method for A/B testing.
In my (again humble) understanding the benefit of doing this in the Bayesian way is both that you can actually get an understandable answer, but also that the answer does not have to be the final result you can continue by adding a loss function for example.
For Bayesian ML see https://probml.github.io/pml-book/book1.html and https://www.bishopbook.com/
For Bayesian statistics see http://www.stat.columbia.edu/~gelman/book/
One of the strengths of Bayesian methods is that you can use a less informative prior to model your skepticism of prior work.
[1]: https://www.sciencedirect.com/topics/medicine-and-dentistry/....
I wouldn't be surprised if many people using frequentist methods have no idea what assumptions need to be true for the methods to be valid. In a Bayesian framework you can't avoid the assumptions