The decline effect and the scientific method (2010)
newyorker.com
newyorker.com
http://www.lastwordonnothing.com/2012/11/05/jonah-lehrer-nat...
If there’s a lesson here, it’s about a widespread human failing. Most people would rather some other clever person distill down all the complex details into a good story for them, preferably in excellent prose. But those distilled stories should never be treated as a substitute for original research results. If anyone really wants ‘the truth’, they’re going to have to slog through an awful lot of turgid and arcane original research and draw their own conclusion.
The utility of the scientific method is that it works even when run by self-interested, flawed, irrational humans. Technology moves ahead. People live rather than die, and become far wealthier than their ancestors. That this process kicked into high gear sometime around the time and place that the scientific method was finally formalized and that formalism successfully popularized, after thousands of years of ragged, slow, and erratic progress, does not seem to me to be a coincidence.
Did I miss any reasons?
Very good article by Warren Buffet that touches on the same issue: http://www.tilsonfunds.com/superinvestors.html
[0] http://uncertainty.stat.cmu.edu/ Chapter 12.
In practice, if you fix p=0.05 and increase your n, the probability that you will find a statistically significant result often increases because your power increases, and in many situations, the probability that the null hypothesis is true is close to zero. (Andrew Gelman uses the example of asking whether there are significant differences between voting patterns of men and women.)
On the other hand, effect size estimates become more accurate as the sample size increases. This mitigates the above issue, provided you actually report your effect size. It also means that small sample studies that report statistically significant results are more likely to overestimate their effect size, which is especially problematic if you are applying null hypothesis significance testing when you know the null hypothesis is false.
For almost all A/B tests, A is actually not much different from B. But due to small sample sizes you will see after the first week, at random, a rather strong positive or negative result. And now the publication bias/selection bias kicks in. If you see strong negative results in the first week you will quickly give up and start a new A/B test. If the initial results show positive results you get excited and keep testing but then in most circumstances you get reversal to the mean at high sample sizes. This would most likely also have happened to the experiments you terminated early but you selected them away and in your memories only for the good initial results a reversal to the mean often seem to happen.
Most A/B tests, ran for a long enough time, will show insignificantly differences. Blogs may give a different impression but again explained by publication bias.
I always compare A/B testing to genetic mutations. Almost none have a strong impact on the fitness of an animal but once in a very long while you have a positive one. Luckily they accumulate and you can get some impressive results with A/B testing (aka natural selection)