Science fails to face the shortcomings of statistics
sciencenews.org
sciencenews.org
One other big error I'm reasonably confident about (though I'd welcome corrections from those who know more stats) is that the p-value, in addition to its other faults and misinterpretations, is usually used in a manner that assumes a Gaussian distribution. While the Central Limit Theorum does mean we tend to see that more often than some other distributions, it is not true that it is safe to simply assume your data is Gaussian. You really need to demonstrate that it is, first, then you can start using Gaussian-based tools on the data.
A typical controlled science experiment is designed to take measurements of multiple groups where one variable is different between the groups and others are controlled. We wish to see if the variable of interest has an effect.
Therefore the commonest statistical test is to determine if the mean value of the groups are different (t-test for two groups, anova for multiple groups) e.g. does mean blood pressure increase when on a high salt diet.
If I understand the CLT (big if) then the distribution of the _mean_ of a sample is, by the CLT, going to be Gaussian, regardless of the distribution from which the actual measurements are drawn. i.e. for comparing group means it it doesn't matter if my data is sampled from Guassians or not.
Of course that leads to the question of if a significant difference in group means is really relevant in a given context.
Yes, for an infinite number of samples. The rate at which is converges to a Gaussian, though, is strongly dependent on the distribution from which the measurements are drawn.
The Central Limit Theorem applies only to samples which have finite mean and variance (see: http://en.wikipedia.org/wiki/Central_limit_theorem).
Take a distribution which has infinite variance or mean and you can wind-up instead with one of the fractal distributions which Mandlebrot studied.
see http://en.wikipedia.org/wiki/Stable_distribution
Distributions with infinite variance or mean are more common than one might imagine. Some might argue the stock market would qualify.
But moreover, in most situations, the other variables cannot be absolutely controlled but rather are only roughly controlled. How roughly depends on the field and is often the rub.
Also it is bias in uncontrolled variables that is most concerning. (More than uncontrolled noise although that will reduce the power to detect real differences).
"There is increasing concern," declared epidemiologist John Ioannidis in a highly cited 2005 paper in PLoS Medicine, "that in modern research, false findings may be the majority or even the vast majority of published research claims."
That said, I agree with the conclusion that published research is full of incorrect results. I basically don't believe anything that comes out of a single paper unless it's been confirmed by unrelated work. That's the fault not just of incorrect statistics but of many other things, not the least of which is that doing research is hard and there is no way of knowing your end result is right.
http://lesswrong.com/lw/1gc/frequentist_statistics_are_frequ...
Suggestion b is both radical and very thought provoking. Which, at this point, I expect from Yudkowsky.
That's certainly true. The only thing you can say here about statistics is that it's hard, much harder than it looks.
Popular media reports with phrases like "Over 80% of..." create the impression that statistics can be encapsulated in a simple number, but that's far, far from true.
More formally, most attempts to "lift" statistical inference into logical inference are unsound. A lot of scientists have, at least informally, a mixed model of: we acquire facts statistically (via e.g. hypothesis tests), then once we have those facts, we reason about them logically (using the usual rules of logical argument). But if you try to formalize this, e.g. building a system that derives facts from data by statistical hypothesis testing, and then uses first-order logical inference on the resulting fact base, you quickly get paradoxes.
This is important, but not quite accurate. It would be more correct to say that this is what the P value means if the model is correctly specified. And in the social sciences (or, for that matter, biology), this is almost never the case.
When I was an undergrad taking econometrics, this was incredibly frustrating. I swore there had to be something I just wasn't getting; why did scientists put so much credence in numbers that rely on assumptions that they know to be false? Of course, at the same time, I love microeconomic theory, which lies on a similarly fictitious basis.
Over time, I relaxed a bit in my attitude toward statistics. While I don't mean to diminish the importance of proper, rigorous methodology, the fact is that statistical methods are just a narrative device. They give us a way of telling plausible stories and discarding implausible ones. We'd be foolish to believe that we can always tell correlation, causality and coincidence apart, but we do a better job by using statistics than we would without.
Indeed. I had a similar realization when I observed that the estimated parameter error on a chi-square fit does not depend on the actual chi-square value itself. This seemed preposterous to me, shouldn't the parameters be more uncertain if the fit is bad? Then I came across this passage in Numerical Recipies that said something like "remember that all of this is under the assumption that the model being fit to is actually the one from which the data points are drawn. If the reduced chi-square value is >>1, then that indicates that this is not the case and then the entire procedure is suspect."
Recall that 95% significance means "5% chance this correlation is a fluke". Correlating 20 questions with 20 others gives you 400 possibilities; they found 24 correlations. Given 400 trials, if there was no real correlation, they should've found about 20 flukes.
Needless to say, I was overall unimpressed by their results.
If you have only studied first year statistics this is what you will learn as "science". As soon as you get to subsequent study, you realise why hypothesis testing is almost always the wrong approach.
I'm not sure who this guy has been talking to, but it sure hasn't been statisticians!
I totally agree that translation of results usually makes the jump from 'statistically significant at a p < .5' to 'fact' - and that's where the problem is.
Scientists are hands-on technicians, theoreticians/statisticians, fundraisers, and team leaders/project managers. Four distinct skillsets.
They are usually selected bassed on their ability to do bookwormish things such as ace tests in their teens and early twenties.
The profession goes against the basic principle of capitalism - division of labour.
I'm kinda amazed anything gets done by scientists, and this article is validating my intuition.
And I think larger labs do frequently separate out a few of these. I'd love to have a better understanding of how industry labs differ from academic labs in that respect.
Most of the useful scientific results synthesize previous results from various fields. There are very few observations incoming that actually add to or question the prevailing theories of a given field. It seems to me from my short experience that most of the work is in reducing theories, and experimentally confirming the reductions. To be useful as a scientist, there's a good change that you'll need to be an expert in at least two things.
I don't think it should be surprising that generating novel knowledge requires substantially more overhead than does acquiring and exploiting it (eg. engineering, or any practical pursuit where division of labour works so well). This may not apply to the medical research that was mentioned in the article, (the system is too complex for theoretic reduction to be useful) but I doubt that physical, ecological, mathematical, etc. sciences would progress as quickly without their practitioners having broad knowledge of their subjects.