Theory-testing in psychology and physics: a methodological paradox (1967) [pdf]
fisme.science.uu.nl
fisme.science.uu.nl
Meanwhile our eager-beaver researcher, undismayed by logic-of-science considerations and relying blissfully on the “exactitude” of modern statistical hypothesis-testing, has produced a long publication list and been promoted to a full professorship. In terms of his contribution to the enduring body of psychological knowledge, he has done hardly anything. His true position is that of a potent-but-sterile intellectual rake, who leaves in his merry path a long train of ravished maidens but no viable scientific offspring.
The problem boils down to the insistence on statistical significance testing. Whenever you test an intervention (priming experiments or whatever a psychologist might look for), it's reasonable to expect there's some effect -- even if it's negligibly small or practically useless. But significance testing is about proving that the effect is not zero, so you will always get a significant result as long as you collect enough data.
This is why I am so frustrated when I see a psychology paper claiming something like "Women wear pink shirts while more fertile" and the paper says the effect is statistically significant, but doesn't say how big it is. Maybe it's a completely unimportant effect.
This is why someone (I forget who) suggested that null hypothesis significance testing (NHST) should be renamed Statistical Hypothesis Inference Testing, for the aptness of the resulting acronym.
There's a whole legion of other common problems (enough to write a book on [see my profile]), but in psychology especially this is a big one.