Statistics Done Wrong – The woefully complete guide
refsmmat.com
refsmmat.com
I'm currently working on expanding the guide to book length, and considering options for publication (self-publishing, commercial publishers, etc.). It seems like a broad spectrum of people find it useful. I'd appreciate any suggestions from the HN crowd.
(A few folks have already emailed me with tips and suggestions. Thanks!)
(Also, I'm sure glad I added that email signup a couple weeks ago)
People also need to realize what while the Discussion and Conclusion section of publications may often read like statements of truth, they're usually just a huge lump of spinoff hypotheses in prose form. Despite my frequent frustrations with the ways science could be better, the overall arrow of progress points in the right direction. Science isn't a process where the goal is to ensure that 100% of what gets published is correct, but whereby previous assertions can be refuted and corrected.
Edit:
To be more specific, I think the statement in your Introduction is overly critical: "The problem isn’t fraud but poor statistical education – poor enough that some scientists conclude that most published research findings are probably false". I would change it to say: "conclude that most published research findings contain (significant) errors", or something along those lines.
http://www.plosmedicine.org/article/info:doi/10.1371/journal...
He's drawn some criticism for the paper, and perhaps things aren't as bad as he makes it seem, but it is true that someone has suggested most findings are false.
I may tone down the Introduction slightly.
Also regarding John Ioannidis's essay (not paper):
First, he uses the blanket term "research" in his meta-analysis (or at least examples) but his work seems focused primarily on medical research studies. Second, I'm not sure he clearly defines what it means to be "False", or for "most" published research to be "false".
Let's say there is clearly a right and a wrong answer to a question, and up until yesterday, publications A, B and C had concluded the wrong answer. But someone releases a newer, more rigorous finding D that refutes A, B and C conclusively and choses the correct answer. I wouldn't consider this particular field to be 75% wrong after the publication of D. (Though it accurately could have been described as close to 0% conclusive before D). For any particular line of inquiry, the quality of research in this area seems like it should be shifted strongly toward the maximum exemplar of this body of work, and not it's average.
As a scientist, I think this is probably correct. In my experience, a great majority of publications draw improper statistical conclusions, and I believe many of these are wrong in substance.
"What if the broad strokes are right but the statistics are sloppy?"
Publishing statements as statements of truth when they are improperly or falsely backed up would be better described as politics than science.
Nope, instead it's not statistically significant, so there's no reason to think it wasn't just random chance that made it 'close' to statistically significant.
But if your results are statistically validated, but only because you used statistics incorrectly, you've crossed the line to broadly-true-but-with-sloppiness?
Nope.
There is room for qualitative research. I like qualitative research. But qualitative research has, righty, a different sort of impact.
If you are claiming your research is quantitative, then you need to live and die by good statistics. That's the implicit promise of providing the statistics in the first place, otherwise why do statistical calculations at all?
Or what if one of the experiments in a paper is well-designed and well-executed and supports a hypothesis with very high certainty but some of the other experiments were sloppy or botched? Should the conclusions of the entire paper be labeled as "wrong"?
I find it pretty funny when critiques on statistical rigor in science arrive at language with words such as "mostly" and "wrong".
This is only a good working assumption of some (open access) journals and of papers (co-)authored exclusively by nationals of some countries. That's a lot of papers.
The main conclusions? One or two minor side points? What if the broad strokes are right but the statistics are sloppy?
If the main conclusions are right but the statistics are sloppy the paper is true, not false.
My confidence in what Ioannidis published went up significantly on learning that epidemiology is mostly bullshit[0] and "Bayer halts nearly two-thirds of its target-validation projects because in-house experimental findings fail to match up with published literature claims, finds a first-of-a-kind analysis on data irreproducibility."[1]
I hope the author of the textbook does not listen to you.
[0]http://lesswrong.com/lw/72f/why_epidemiology_will_not_correc...
[1]http://blogs.nature.com/news/2011/09/reliability_of_new_drug...
I'm not arguing against the work itself, or against more rigorous application of statistics. I'm just arguing against sensationalistic and inflammatory language. Anyone who practices science in a particular field for any length of time will have a pretty good idea of what work is "good" and "bad". Certainly they are smart enough to ignore previous work that has been refuted and/or retracted, and it's not really fair for this previous work to contribute to assessments of what % of the field is "wrong".
One thing that is missing is like a Summary/Checklist chapter that tells you what you SHOULD do, in a few common scenarios to avoid all the mistakes presented in the previous chapter. I know its not that simple, and it depends a lot on how you are testing and what you're actually trying to achieve, but a few examples wouldn't hurt.
For example: I have two sets of measurements, and I want to know: * is there a statistically significant difference between them * if yes, how much is the difference
A somewhat simplistic way I'd do that is to do a two-sample t-test for question #1, and to compute statistics for the difference between the samples (mean, median, confidence intervals), but doing just that I might've already committed some of the mistakes that your site warns about, for example I completely disregarded the power of the test.
FWIW I like parts of this book on statistics which focuses on statistics in the domain of computer systems / network, although it is rather too long: http://perfeval.epfl.ch/
Sphinx can output Latex for printing to PDF. We'd just need the source.
BTW, excellent refresher for statistical screw-ups, I had forgotten a quarter of these (and never learned the rest.)
I don't accept donations, but sign up with your email and I'll let you know when you can pay me for a copy.
I'm writing using reStructuredText and Sphinx because they give me a bibliography, an index, full-text search, cross-referencing, custom environments (e.g. boxes for examples, tips, etc.), and all sorts of other cool features. For example, if I needed matplotlib plots, I can embed the code in my documents and Sphinx will generate them.
I'd immediately put the book on Leanpub if I could get some of these features. A Leanpub using a customized Sphinx would be incredibly useful for people writing textbooks.
There are a million other UX tips that you can probably get from a real expert, but the black one I noticed.
... might be nice to add to your list.
I would donate to that goal.
Arguably this would be a greater benefit to humanity than all the millions poured charitably into cancer research etc.
Of course, medicine might be unique as a domain in which individuals are willing to pay vast sums of money to obtain slightly more trustworthy research conclusions, and the profit motive has obvious conflicts with "benefit to humanity" (if someone pays you to research a treatment for their disease, do you post the findings when done? Or hold them privately for the next person with the same problem?). But maybe there are other domains in which the market could support a (non-billionaire's) project for better-validated research.
[1] https://www.scienceexchange.com/reproducibility
[2] http://www.newscientist.com/article/mg21528826.000-is-medica...
They perform meta analysis of studies and talk about the validity of their statistical methods. http://summaries.cochrane.org/
Read about it in the book Bad Science by Ben Goldacre.
Notably, this is easier in computer science as you don't need to wait for hundreds of patients to turn up having a certain condition.
Learning the right way takes a lot of work, there's a lot of ways to analyse things, each one wrong/right in different situations. (Even teaching something as "simple" as the correct interpretation of a p-value is hard.)
If you are interested in the difference of a metric scaled quantity between two groups do the following:
1.) Add 4-5 plausible control variables that you do not document in advance (questionaire, sex, age...).
2.) Write a r-script that helps you do the following: Whenever you have tested a person increment your dataset with the persons result and run a:
t-test
u-test
ordinal logistic regression over some possible bucket combinations.
3.) Do this over all permutations of the control variables. Have the script ring a loud bell when significance is achieved so data collection is stopped immediately. An added bonus is that you will likely get a significant result with a small n which enables you to do a reversed power analysis.
Now you can report that your theoretical research implied a strong effect size so you choose an appropriate small n which, as expected, yielded a significant result ;)
[1] http://euri.ca/2012/youre-probably-polluting-your-statistics...
Or maybe I'm projecting?
Not true. It's all about risk vs. payoff. Some things are low risk enough we can go by gut, others we need more evidence. It's all about tuning for false positives and negatives.
EDIT: Added quote of what I was responding to.
http://opim.wharton.upenn.edu/~uws/
have published a lot of interesting papers advising psychology researchers how to avoid statistical errors (and also how to detect statistical errors, up to and including fraud, by using statistical techniques on published data).
Other ways of doing this include JAGS (http://mcmc-jags.sourceforge.net/) and Stan (http://mc-stan.org/)
The advantage of statistical modelling is that it makes your assumptions very explicit, and there is more of an emphasis on effect size estimation and less on reaching arbitrary significance thresholds.
But despite this, statistics done well are very powerful.
The irony of people who use the "damned lies and statistics" quote snidely is that the "statistics" part is not referring to the field Statistics but the plural version of the noun statistics, which of course are easily abused. The field of Statistics is all about NOT abusing statistics.
Yeah, it always makes me wonder when they have to put the "science" in the name. "Computation" seems so much more timeless and elegant than "computer science", for instance. It's almost like "Democratic Republic" for nations.
At this point, if someone published a study stating that we needed to eat not to die, I'd be skeptical of it.
Schoenfeld, J. D., & Ioannidis, J. P. A. (2013). Is everything we eat associated with cancer? A systematic cookbook review. American Journal of Clinical Nutrition, 97(1), 127–134. doi:10.3945/ajcn.112.047142
They did a review of cookbook ingredients and found that most of them had studies showing they increased your risk of cancer, while also having studies showing they decrease your risk of cancer.
I think bacon was a notable exception -- everyone agreed that it increases your cancer risk.
http://en.wikipedia.org/wiki/Oil_drop_experiment#Fraud_alleg...