It is time to stop teaching frequentism to non-statisticians (2012)
arxiv.org
arxiv.org
- Pre-register all studies, declaring sample sizes and power analysis.
- Report results regardless of outcome. Eliminate the "we only publish stat sig results" baloney.
- Report confidence/credible intervals, adjusting for multiple comparisons as appropriate. Plot the posterior distribution of the effect size if appropriate.
- Publish all data and code.
- Provide funding for duplicating important studies.
I'd add one additional technique: Specification curve analyses. Bonferroni etc. alone won't help against systematic bias in a field and/or misconduct, and specification curves are easy to do, e.g. with specR [1].
[edit] you can get posteriors with bootstrapping etc. - it's not precisely the same but better than nothing.
https://slimemoldtimemold.com/2022/07/21/on-the-hunt-for-gin...
Bayesian and Frequentist methods are not nearly as at odds as posts like this suggest. Frequentism is mostly about how methods should be evaluated. Bayesianism is mostly about how to incorporate different sources of information. You can assess the Frequentist properties of Bayesian methods!
Larry Wasserman wrote a great post about this topic here: https://normaldeviate.wordpress.com/2012/11/17/what-is-bayes...
I always report confidence intervals front and center, and bring in point estimates and p values as supporting characters. And of course discussing to what extent the study design supports causal conclusions.
It's like a programming language - there is no "best" one in general, you just have to learn how to choose the best tool for the job at hand, understanding each tool's strengths and weaknesses. It's not that mysterious.
The biggest issue with frequentism is the assumptions. I can't rattle them off like I used to be able to, but almost every real world scenario where statistics are useful are going to violate some of them, and yet frequentists will simply carry on.
It's really a cultural issue around 'correctness', and frequentism is often reduced to an appeal to authority.
> P-values can and are used to prove anything and everything. The sole limitation is the imagination of the researcher. Fleeting exposure to a 72 × 45 pixel image of the American flag turns one into a Republican; walking through a door (an “event segmentation”) damages your memory; Keynesian theory is wrong; Keynesian theory is right; selenium causes cancer; selenium cures cancer. This list could (and will) go on in perpetuity.
Sadly true and well written. However, the author is over certain that using Bayesian statistics will make things better.
This is more of an argument not to teach anyone statistics at all, other than some unspecified elite that the author feels qualified to learn it.
By the way, the author has some interesting obsessions:
https://twitter.com/FamedCelebrity/status/155513494784911360...
https://twitter.com/FamedCelebrity/status/155516247586318336...
https://twitter.com/JohnZmirak/status/1554969385860796420?s=... (retweet)
"How many students and teachers noticed that all of the statements were wrong? As Figure 1 shows, none of the students did. Every student endorsed one or more of the illusions about the meaning of a p-value. One might think that these students lack the right genes for statistical thinking and are stubbornly resistant to education. A glance at the performance of their teachers, however, indicates that wishful thinking might not be entirely their fault. Ninety percent of the professors and lecturers also had illusions, a proportion almost as high as among their students. Most surprisingly, 80% of the statistics teachers shared illusions with their students."
[1] http://library.mpib-berlin.mpg.de/ft/gg/GG_Null_2004.pdf
The problem with frequentism and the null ritual is that it makes statistics easier (just do this one test, read this one number) and renders some kinds of mistakes somewhat harder (publishing a single false positive).
At the same time, it makes some bad mistakes much easier, most notably ignoring power, false positives, the garden of forking paths, type S errors and publication bias.
The inherent problem is that the null model is not what people assume it is, and the method (or at least the established canon of approaches) don't make you think about it.
If you use bayesian methods, you're pretty much forced to spend more time considering the effect size and credibility of your results, and you're basically required to report them.
This means that even non-competent bayesians probably have a better contribution to cumulative science.
Frequentist statistics does not force you to accept any hypothesis test with a result p < 0.05 as definitive proof of something. It does not forbid considering prior probability of a result. It just doesn't formalize the consideration of prior probabilities because it is hard to distill this consideration into a formal recipe.
Everyone agrees that Bayes' Rule is valid and important, the question is when and how best to use it.
> If you use bayesian methods, you're pretty much forced to spend more time considering the effect size and credibility of your results, and you're basically required to report them.
You're not forced to do those things well. Any scientific method can be cargo culted.
The problem that I care about is not whether frequentist statistics can be taught and used well. They can, and I try in my teaching to do so.
The problem is that empirically, frequentist statistics is a fig leaf for a ton of extremely problematic work. And pushing bayesian thinking is currently our best chance to fix this, because it's easier to do a shift in the mental framework than to fix the perception of an existing framework.
Bayesian methodology benefits from having relatively much more statistically sophisticated practitioners, which leads to an optimism bias when we imagine how it would scale up.
I agree w/ the parent poster than the fundamental issue is probability: if you are talking w/ people w/o background in stats, you will have a really hard time to go beyond a true/false statement.
Moreover, one of the most effective (in $ terms) application of statistics in recent times is A/B testing. While you can do "Bayesian A/B testing", the basic methodology is fundamentally frequentist. Mistakes there can be hedged through better tooling / UX (to avoid peeking, etc.), as effectively as using Bayesian statistics.
Ok, that explains why people bother spend so much energy badmouthing frequentism, but it's a very bad framing anyway, bothering a lie. Frequentism is not the null ritual. In fact, it's almost completely compatible with bayesianism, the one large difference being the freedom to set priors before doing your analysis.
If the article was titled "It's time to stop teaching the null ritual to scientists", nobody would even disagree.
This may be true of "frequentism" in a strict sense, but many statistical methods in common use that are often described as such (including, arguably, NHST) are not consistent with reasonable versions of the likelihood principle https://en.wikipedia.org/wiki/Likelihood_principle . In a sense, one might feasibly argue that these methods are not even properly frequentist.
There will always be no instances where individuals err in the way they apply statistical methods. But that doesn’t mean that there is no value in moving the “default” in statistics to a more intuitive methodology from one that is so obtuse that it’s common for relatively advanced practitioners to stumble over it
It's possible that it can happen the other way around as well, but my impression is that it happens less often.
> https://twitter.com/FamedCelebrity/status/155513494784911360...
> https://twitter.com/FamedCelebrity/status/155516247586318336...
> https://twitter.com/JohnZmirak/status/1554969385860796420?s=... (retweet)
None of those twitter threads have anything to do with statistics pedagogy at all. What sort of irrelevant, ad hominem argument are trying to make here?
Alas, the article is badly written and unprofessionally laid-out, and, ironically given the title, isn't written with a non-statistician in mind.
It's not any more thoughtful, it's just the opposite.
If you are talking about his statistics opinions, I don't understand enough statistics to know if it's right or wrong or not even wrong. But his other claims gives him a substantial respect-boost regardless of whether he's right or wrong. I simply like contrarians more.
The yellow flag in those tweets is the hyperventilating culture warrior attitude. It speaks poorly of ability to humbly and dispassionately analyze topics, regardless of which side it comes from.
False. It's a pretty reliable proxy if the Establishment is known or plausibly suspected to benefit from falsehood and\or misrepresent truth. If you have a reliable predictor of falsehood, then something which always opposes it is a reliable predictor of truth. I don't know if the statistics Establishment is one such example (probably not), but the trans-activism and war-media Establishments certainly Are reliable predictors of falsehood, and have been shown to be on multiple occasions.
That's why I like the author's opinions on those topics more than on statistics, even if I'm not necessarily certain of their truth. He's going in the rough general direction of truth.
>hyperventilating culture warrior attitude. It speaks poorly of ability to humbly and dispassionately analyze topics, regardless of which side it comes from.
False Again. Plenty of smart people have this attitude, the notorious example off the top of my head is Nassim Nicholas Taleb, incredibly arrogant and quick to ungraciously fail into combat under criticism - yes - but do you actually want to take him one-on-one in statistics? For a more distant example, see Newton's hilarious pettiness in dealing with rivals. Your heuristic yellow flag would have lead you to dismiss him faster than you can say "Universal Gravitation", and we both know you would be utterly and irredeemably wrong.
Easy to misunderstand. NHST creates false beliefs, and that's even when you do everything correctly. If you accept NHST, you believe in "psi", because Bem has proven it.
It reminds me when I learned about Tyndale translating the Latin Bible into English. There was a pushback in creating a commoner’s version of the Bible where the uneducated would be able to read the Bible for themselves. The Catholic Church was later involved in his death.
Just because a text is not pandering to the lowest common denominator with respect to literacy, it's not "full of [itself]".
It has the form of a scholarly article, but the author has no credentials, no institution, doesn't cite any references and to my knowledge has not submitted the article for any kind of peer review.
This is effectively Some Dude On Twitter but on Arxiv.
I find the greatest proponents of Bayesian thinking to be in the rationalist community. So to get across just how wrong everyone is about just about everything pretty much always consider that the pejorative lessons on predictable irrationality and cognitive bias are delusions that come about through the flaws of statistical reasoning. From first principles we know that superhuman chess AI are also predictably irrational and cognitively biased, but that this isn't actually irrational or biased so as to desire us to abandon the choice. Even in chess - much simpler than reality - the correct solution was too large to fit within our universe.
It is difficult to articulate the magnitude of this humility. We couldn't fit the chess solution into our universe, but how complex is chess relative to other problems? If the complexity of chess was a grain of sand or even a single atom than it would be an understatement to claim that the more general problems are only as large as our universe.
Oh. How we all wish the word civilian here din't mean scientists in every field publishing in every journal.
And how I wish my reviewers didn't force people to do this :(
Civilian is someone who's never been in the armed forced and therefore ignorant of how things really work in the military and in battle.
So it's basically a metaphor for a someone who's never studied statistics in depth and therefore ignorant of how statistic really work.
That said, every step that forces everybody to acknowledge uncertainty is good. Given our progress e.g. in getting people to recognize how important power is (e.g. since Cohen, 1992 at the latest), I'm very pessimistic.
But while we may not be able to get everyone on board the bayes train, we can at least force people to openly show that they don't care about good statistics.
Hoeffding's inequality, Chebyshev's inequality, and Chernoff's inequality are broadly applicable, and are therefore less likely to be misused. They also don't require philosophical assumptions about subjective probabilities.
In the theory of Multi-Armed Bandit algorithms, compare the Bayesian approach (Thompson sampling) with the Concentration Inequality approach (the UCB algorithm). The former assumes that payoffs are in the 2-element set {0,1}, while the latter allows payoffs to be in the interval [0,1].
I think Bayesian is cool, but it imposes a big burden with choosing sensible priors.
Heh. So people noticed.
'The "Null Ritual"', Marc Green https://www.visualexpert.com/Resources/nullritual.html
Is a linear model 'better' than a regression developed using a tree method allowed to run over hundreds of parameters?
The amount of time I've been forced to take explaining random forest, and why we can trust it's R2 more than we can a linear regression is ridiculous.
They are definitely useful as part of bigger models that we cannot explicitly express them in multiple fields. Physics, chemistry, ML, they all utilize statistics to make useful models and that is great.
The issue is that decision making is a singular point, and statistics are very good at not taking any responsibility for their bad predictions.
If you make a decision with an expected outcome the statistician will come out swinging. If it was the unexpected outcome then you are faced with “Well you were unlucky, that’s statistics”.