How to Hang Yourself with Statistics
math-blog.com
math-blog.com
Keep in mind that people flip coins and get five heads (or five tails) in a row all the time. With a p value of only five percent, one in twenty published papers reporting a p value of five percent will be wrong purely by chance.
That's only true if 50% of the hypotheses you test are true, and your experiment is so good that you have no false negatives. In typical medical trials, on the other hand, sample sizes are small enough that there's perhaps a 50% chance of false negative. (If you see a difference between groups with small sample sizes, you can't tell whether it's due to the tested medication or just chance, so you conclude there's no statistically significant difference.)
Under realistic circumstances, the odds of a p < 0.05 result being true can be as small as 45%.
I've written more on this problem here: http://www.refsmmat.com/statistics/#the-p-value-and-the-base...
The other point I particularly like is the discussion on mean values and robust statistics. Most analytics packages just report means, losing so much information in the process. Again I bet many bad decisions are made due to poor tools.
[1] http://www.plosmedicine.org/article/info:doi/10.1371/journal...
That said, it's worth being reminded occasionally that statistical significance is in many ways just the first step on the road to knowledge, not the last. But when it comes to false findings, stuff like false negatives and exploratory analysis are far more impactful.
But p=10^-6 doesn't mean, as commonly believed, that there's only a one-in-a-million chance that the proposed hypothesis is really false, nor does it even mean what many more-statistically-savvy people think it means, that if the proposed hypothesis were false, there would only be a one-in-a-million chance of observing test data as extreme as what was observed. No, what it really means is that – and here's the part most people miss – assuming that the researchers' model of the underlying data-generating process is correct, then, if the proposed hypothesis were false, there would be only a one-in-a-million chance of observing test data as extreme as what was observed.
Yes, as the p-value becomes smaller, it does indeed become easier to believe that the hypothesis of interest is true, assuming that the humans didn't screw up the model. But, in any complex work, I'm going to have a hard time believing, sans replication, that there's not a reasonable chance of humans screwing up.
To me, then, p=10^-6 is the new p=10^-2.
EDIT: Replaced Unicode superscripts (10⁻⁶) with circumflex notation (10^-6) because the superscripts weren't showing up on my Nexus 7.
It is a difficult task for many people, who have been taught facts for decades, to accept that objective knowledge is hard to come by. But everyone understands the value and properties of a crude model.
So I agree with you. If it were up to me, we'd all report evidence intensity in decibels, anyway.
http://math-blog.com/mathematics-books/
but the subtopic with the most disappointing recommendations was actually the subtopic on statistics. Two of my favorite online articles on statistics education
http://statland.org/MAAFIXED.PDF
and
http://escholarship.org/uc/item/6hb3k0nz
both point to better books on statistics and the key issues in the discipline.
I currently have Feller (based on reading www.ams.org/notices/200510/comm-fowler.pdf) but don't have a good "taste" for what would be the best books for a self-study approach to statistics. There is also the "Teaching Statistics: A Bag of Tricks" book and I was considering dropping it into the mix.
I am probably over-thinking the book choice and would be fine to just dive in to anything, but if I would prefer to not pick up references that are going to drive me off a good path to start . . .
The submitter is the owner of the blog.