But to be fair, we should take things on their merits. Contrasting the two papers, Bem's paper sticks very closely to the standard form and 'rules of engagement' (if you will) of published academic research. This counter paper on the other hand, has only the surface appearance of proper scientific method. For example, in making one if its main points it refers to a hypothetical casino where Bem could've made an 'infinite' amount of money, and to the one million dollar Randi prize, neither of which reach the standard of proper experimentalist scientific evidence. It does this by way of justification to one of their main points which is that because 'extraordinary claims require extraordinary evidence' we should be able to set the prior expectation of H1 - in their words 'for illustrative purposes' - to .00000000000000000001, which they then go on to demonstrate, makes the results non-significant.
But where does the 0.00000000000000000001 come from? It could, just as easily, be twice or half that figure. Thus not falsifiable and therefore not justifiable as an extra bit of arithmetic that Bem's paper must qualify.
To put this in terms this audience will understand, that's kind of like saying well no, because you are using Java and 'everybody knows' Java is slow I think we should multiply your benchmark figures by, oh, let's say, one half. And then we see that java throughput is quite poor, as expected. In fact, we should multiply all java benchmarks by some number like a half but I'm not going to be specific about it because actually its pretty much just an arbitrary number I made up. So, reading past the reference to Bayes and some nice formulas that's just arithmetic in my book.
Note, I'm not saying that there are no flaws in the Bem paper. Everybody can see that it's very likely they'll be something wrong with it (and the desk drawer problem to which the above paper also makes extensive evidence is a likely though not conclusive contender) but I think its only reasonable to hold up the criticism's to the same standard as what they are criticizing. Perhaps that way you'll be more likely to find the actual truth of the matter.
That's kind of the problem: published academic research is broken in a lot of ways. If you haven't read it yet, the New Yorker's recent article on the decline effect touches on several of the reasons why: http://www.newyorker.com/reporting/2010/12/13/101213fa_fact_... And some prominent medical journals started fighting against publication bias a few years ago: http://www.smh.com.au/articles/2004/09/09/1094530773888.html
I think you're holding this response to a much higher burden of proof than it needs to meet to be a proper refutation of Bem's claims. You're right that the response doesn't appear to use the "proper scientific method". But that's because it doesn't use it at all, and it doesn't need to. There's no hypothesis to test and no experiment to run in order to point out flaws in a paper that does claim to be the result of the scientific method.
It makes the very good point that this paper makes lots of statistical tests, and then bases big claims on the small minority that showed a significant effect. This is in no way restricted to psychology, drugs companies for example do this all the time. It's cheating, whether you realize it or not.
Statistical significance only tells you that a result is unlikely to be a fluke -- not that it definitely isn't a fluke -- but the more tests you do, the sooner you'll see a fluke on average.
In other words, if you toss a coin 1,000 times, then it's hideously unlikely that you'll see a run of 100 consecutive heads. But if you toss the coin 100,000,000 times, you shouldn't be too surprised to see that 100-toss run buried in there somewhere, even though the odds of getting 100 in a row are so small.
Right?
The only way to find out is to do enough flips to eliminate the chances of your final result being influenced by statistical flukes. Measuring small differences, like trying to answer "does a coin preferentially land on one side vs the other?" usually takes hundreds of thousands of tests to guarantee you're seeing objective data, rather than seeing patterns in noise.
the very moment you peek, your data is tainted from future testing.
If you toss 100,000,000 == 2 ^ 27, you should only expect around 30 in a row. To have a good chance of getting 100 in a row, you need about a billion squared times more.
And the problem is MUCH worse than described above: Let's say you test 1000 wrong hypothesises with p=0.05; 50 of those will be accepted as true, even though all are wrong. If you test 980 wrong hypothesises and 20 right ones, more than half of those that pass the p=0.05 "golden" significance test will in fact be wrong.
Now, when you see a medical journal with 20 articles using p=0.05, which do you think is more probable - that 19 are right and one is wrong, or 19 are wrong and one is right? The latter has a much higher likelihood.
The whole field of systematic reviews and meta-analyses has developed around the need to aggregate results from multiple studies of the same disease or treatment, because you can't just trust one isolated result -- it's probably wrong.
http://en.wikipedia.org/wiki/Meta-analyses
Statisticians working in EBM have developed techniques for detecting the 'file-drawer problem' of unpublished negative studies, and correcting for multiple tests (data-dredging). Other fields have a lot to learn...
Regardless of the true reason, these are never carried out before a new drug or treatment is approved (because there is usually one or two studies supporting said treatment, both positive).
And if you have pointers to techniques developed for/by EBM practitioners, I would be grateful. Being a bayesian guy myself and having spent some time reading Lancet, NEMJ and BMJ papers, I'm so far unimpressed, to say the least.
I would be a lot more relieved if people with the power over other peoples lives grokked Baye's rule and priors.
Bayes's rule doesn't help with the point that suggestive evidence is not convincing evidence. It just points out that prior beliefs are part of the equation, but will hopefully pale in comparison to actual data. In fact, I was taught to set practically useless hyperparameters to ensure that they do. No one does that outside of an experiment.
Let's say I believe (I don't) the height of pygmies is normally distributed, where the mean is also normally distributed with mean 130cm and standard deviation 10cm, and the standard deviation is inverse gamma distributed with shape 7cm and scale 1cm. Assuming the height is actually normally distributed with mean 160cm and sd 15cm (it isn't), how many pygmies must I measure to admit that P(height>160cm)>20%? I'm not sure I can even do the math.
Here P=50% for the unknowable accurate model and P=0.13% for the prior model. How does the situation change when my prior is "sufficiently close to 0"?
In the legal case example, maybe some clarity maybe had in considering what does the prior mean. A answer is: say you have to bet a million dollars on whether the person is guilty or not without knowing anything about the person how would you distribute your million dollars between the two events. Yes it is subjective and personal, but it is hardly ever going to be 50:50. One can push the $1,000,000 analogy further. One can fix a cost for a mistake: whats the cost of a wrong conviction and whats the cost for setting a guilty man free. Then the final decision can be based on reducing the financial risk based on the likelihoods.
One may bring the socio-economic status in forming the priors but one may not consider any information source that considers the accused.
Lets take the example. There is a one to one correspondence with fictitious counts and priors. One way of encoding a 50:50 prior is to construct a possibly fictitious but representative past of (say) 2000 samples split into 1000 guilty and a 1000 not-guilty. After each prediction and assuming that the truth gets known one has to update the counts appropriately, so that the next time we use a different prior.
Our initial prior may be wrong but it will approach the correct one asymptotically. But how fast it approaches the true prior depends on how wrong our initial prior was.
It's when the hypothesis predicts a pattern that you haven't noticed yet and the pattern is confirmed by experiment, that's when you know you have something [a theory].