It’s like how you wouldn’t hear people in the US promote imperial measurements for home baking. When you dominate, you don’t have to advocate.
My college offered stats in two tracks: There was a one-semester course for science majors, which was mostly plugging numbers into formulas. The course was utterly baffling for most of the students who took it.
There was also a two-semester course for math majors. It was mainly about proofs, but also had time to go into more depth. You have to know the assumptions underlying a formula, if you're expected to prove it. ;-) But it was only taken by the math majors.
One thing we didn't have when I took stats was computers. I graduated from high school in 1982, and the colleges in my state were just beginning to get computers. I wonder if stats could be approached differently if it could start with nothing but data -- lots of it -- and graphing tools. This could even happen pre-college.
First course was almost purely frequentist
Second was half and half with several weeks approaching the frequentist bayesian approach from both sides
Our third stats course was mostly bayesian though self directed as to which point of view was “most appropriate” for the formulated problem sets. I remember roughly half the cohort used a frequentist perspective on a fairly obvious (but not explicit) bayesian assignment and were severely punished by the course coordinator (fairly for a late year subject tbh)
You can squish it into Bayes by considering it a uniform prior on the real number line that you never update, but you’re not really doing things “in the spirit” of Bayes, then.
IMO, null hypothesis testing is way overused. We should be quantifying the comparison of the null versus alternative hypotheses. People already compare them. Might as well bring some math into it.
If your hypothesis is that A is greater than B, then you're test boils down to "The probability that A is greater than B", which is arrived at through parameter estimation via Bayes' Theorem.
You can consider the distribution of estimates for a parameter to represent a space of possible hypothesis for the true value of that parameter and their absolutely or relative likelihood based on the information available.
The also has the pleasant consequence that hypothesis testing using Bayesian methods nearly always is closer to the actual question someone wants to answer rather than replying with confusing statements such as "the probability of rejecting the null hypothesis", which contains an implicit double negative.
but even in this situation you would need a threshold for decision making, i.e. a p-value.
Bob bob bob'bob bob bob bob bob, bob bob bob bob bob "Bob".
Bayesian decision making typically involves combining a cost function with your posterior distribution.
It's worth pointing out that this is typically not possible in most frequentist frameworks since they work in double negative style assertions and arbitrary threshold rather than modeling the problem itself.
While you can map Bayesian approaches to frequentists tests, this is usually a misunderstanding of Bayesian methods coming from a purely frequentist background. Bayesian analysis is fundamentally more flexible since it just the application of the sum and product rules (you don't even need Bayes' theorem since it can be trivially derived from these) as well as corresponding cost/value functions.
In the most common situations this typically means the expected value well exceeds the expected cost and you're comfortable with the risk if you're wrong.
A prior could be any previous knowledge that you have on the subject, or anything about the world.
Let me give you a silly example. Imagine I do an experiment, on if person A can see the future, and predict coin flips.
Now, imagine that I do this experiment, and lo and behold it passes the 5% threshold P value!
Remember what a P value is. It is simply the percent chance that the results that you got were by chance.
So, if I do that experiment, and it actually gets at the "P value" of 5%, which means that there is a 5% chance of it happening by chance, do you know believe that Person A can see the future?
I wouldn't. I would need a much lower P value to start believing that.
Thats all a prior is. It is using the fact that "Person A can see the future!" is an extraordinary claim. And for that, we require extraordinary evidence. Much better evidence than a mere "5% that it happened by chance". P
P-values are fundamentally Frequentist because this framework see that observed data as random, whereas Bayesian statistics believe our observations are the only thing that is actually known, and all the other parameters are what is random in our experiments. That is: the data is the only part of an experiment that is real, everything else that we can't directly observe is something we must hypothesize about.
Bayesian statistics is sound but I suspect it's often just used to justify biases. It is technically valid to use a prior and de facto never update it, because you know I'll get around to updating my prior next week, or... eventually, cough cough let's be honest, never
It's the frequentist one that gets your biases implicitly, on the form of corrections and hypothesis formulation, so that people don't notice them.
One issue is that Bayes estimates are almost always produced, even if no information is coming from empirical data, and all the information is coming from a prior. So it's possible to produce results heavily influenced by the prior with Bayesian estimation that with frequentist methods would fail completely because of lack of identification of the model, sending a strong signal that something is wrong. This can all be sussed out with Bayesian methods but people often don't do it.
Another more subtle issue is people aren't quite aware of how a prior can deviate from "maximal conservativism". Sometimes, for example, depending on the model, a very flat prior is actually not conservative, and is overweighting tails.
There's other examples too. Basically, yes, Bayesianism forces you to be explicit with your biases, but people are really bad at interpreting the actual impact of those biases in a formal Bayesian framework, or at least, aren't any better at it than with frequentist methods that are available.
If you approach statistical inference from the perspective of accuracy (as the linked paper seems to do) Bayesianism is better to the extent the priors are accurate. This is true a lot of the time empirically, but it does lead to a kind of tautology, in that you're doing the analysis because you don't really know what "truth" is. So if you're right in your priors, Bayesianism is accurate, but then you didn't really need new data as much in the first place; if you're wrong, it's more biased. Basically in the bias-variance tradeoff, Bayesianism makes a bet on reduced variance assuming that the resulting bias will be small enough.
Philosophically, though, there's a completely different argument, which is one of competitive fairness. You might say this doesn't matter, but consider consequential decisions, like hiring or admissions decisions: if someone was making a prediction about you, would you want them to use a strong prior, or something that's maximally conservative and fair?
This philosophy leads to frequentism basically.
My preference is to be maximally conservative in a Bayesian framework, which leads to reference priors, which are often flat in many canonical situations, which is basically frequentism. In other situations you might have a different kind of prior.
To me the linked paper is pretty interesting and makes a good point. On the other hand, I'd rather not make any assumptions about a new result based on past studies on other effects. I'd rather just collect lots of diverse real data and meta-analyze it. There's no substitute for data -- and that includes priors.
I broadly agree with you, but I'm wondering if you would reconsider your qualification as "less complicated" if you consider beginner learners. E.g. someone who knows basic descriptive statistics and probability theory, and is making first contact with inferential statistics. Specifically, assume a learner who knows what an integral is, but is far from proficient with it (UGRAD student, not a GRAD student).
I was reading this paper[1] recently, which highlights two difficulties of teaching Bayesian stats: 1) the mathematical complexity of understanding conditional probability distributions, and 2) the lack of well defined, broadly accepted conventions for what priors to use in specific data analysis scenarios.
I think a computational approach to prob theory could mitigate 1), but 2) remains a problem—the freedom to choose priors, is also a burden...
[1] https://www.stat.purdue.edu/~dsmoore/articles/BayesPedagogy....
Maybe someone here might have suggestions?
The closest thing that comes to mind is "Bayes factors," which has some traction (usage), but apparently they have lots of problems and limitations too, cf. https://www.youtube.com/watch?v=MqeWpR6S4XA
The canonical approach is to build a generative model with a parameter (or multiple for ~anova) that codes for the difference between groups and do inference on that parameter of interest. Most of the recipes taught in statistics classes can be modelled as a regression of some kind (this counts for frequentist stats too, see https://lindeloev.github.io/tests-as-linear/ ). Some advocate to do that inference with bayes factors. Others, like discussed elsewhere in this thread, advocate combining the resulting posterior with a cost/value function, but either way the lesson is that there is less focus on "t-test-vs-anova" because they're the same thing anyways.
I had previously started the BDA course, which is another famous Bayesian course, see https://avehtari.github.io/BDA_course_Aalto/ but I didn't finish it due to travel.
No more excuses in 2024... time to level-up the Bayesian modelling skill ;)
Of course the winners don't rant about anything. But any time we probe the consequences of Frequentist statistics they turn out to be horrific for science, our health, and our planet.
I will do it:
1. Computationally easier
2. often analytical theory available for most use cases so interpretability is high
3. more literature available so you can get unstuck faster if you mess up
4. no accusations of subjective bias in your prior (the con is clear, no ability to leverage subjective expertise)
5. In the asymptotic regime, MLE and bayesian MAP often converge anyways
6. king of hypothesis testing
For most people, it doesn't matter. It matters when you are doing treatment for small sample sizes or other situations that would cause low power.
This advantage can't be understated. Researchers, at least in psychology, almost never seem to care about the magnitude of an effect. It's simply enough to show that some effect happens in some direction. For this (generally acceptable) purpose, frequentist stats are great.
It's not entirely unreasonable, as often extra variables suppress the effect size. E.g. the effect is substantial in people with some genetics, insignificant otherwise. The fact that there is an effect at all makes it interesting as a starting point for further exploration e.g. to determine the mechanism or find the extra factors needed to make the effect significant.
That's not an advantage, that's one of the primary reasons psychology is among the worst subjects hit by the replication crisis.
So not a cheap shot at all IMO.
For example, suppose you are researching the effect of an emotional/negative picture on how a participant makes decision in an economic game. Here, you may think “a big effect size implies this study will have a large practical relevance.” However, following typical statistical designs, an effect size may be largely determined by just how many trials the experiment used for each participant. This is also not the only factor that’s relevant. For instance, it’s debatable how the emotion induced in the task compares to the intensity of real life emotional situations. With these factors in mind, the effect size really doesn’t tell you much that can be applied outside the laboratory.
Maybe in some non-experimental sub-fields (e.g. personality psychology) effect sizes are more meaningful…
At least in my undergrad education, I took two dedicated statistics classes and many more domain science classes that used statistics, all of which were so steeped in the frequentist paradigm that frequentist vs Bayesian wasn't even mentioned. Frequentist statistics simply was statistics for me, until grad school.
It is ultimately the same math, if you are comparing apples to apples. Just different ways of looking at a problem. Sometimes one is better suited than the other.
It’s like as if physicists were arguing over whether Cartesian or Polar coordinates were better. It’s the same damn physics, just expressed differently. In some problems one approach is easier to work with than the other, and can even make seemingly intractable problems solvable. But that doesn’t mean the other approach was “wrong.”
The controversial part is the methodology for statistical analysis built on top of it (e.g., you need a prior but where did that come from?)