I've already posted this link today, but I really recommend this guide: http://www.evanmiller.org/how-not-to-run-an-ab-test.html
Run enough tests, and you will get statistically significant but bogus results.
Many times, when we declare a winner in test and if you keep running it, you may see that eventually it does not perform as good.
But I agree my comment doesn't negate Jabbles' point.
If you run 100 tests, a small number of them will be FALSE positives, their confidence level will be above 95% even though they are, in fact, just statistical anomalies.
That's why it's 95% confidence and not 100% confidence.
What it means, by hiding all your failed tests, is that you are probably only writing about FALSE positives.
If you can't understand that, how can we trust any of your blog posts?
NB/update: Even though I'm generally pretty good at maths, I've always found statistics extremely hard. I totally understand how hard it is because it can produce such mind-boggling counter-intuitive results. But if you don't understand something that's key to your domain you should learn about it, not gloss over it.
Regardless of the statistics though, I personally think http://visualwebsiteoptimizer.com is very good.
However I have seen people here on HN make posts about A/B testing that even with my shallow understanding of statistics makes me raise my eyebrow that they really don't get it.
It became apparent that Taguchi wasn't really appropriate or sufficient for web-based testing, so the team learned, devised, and implemented more appropriate MVT models.[2]
One notable bug that we discovered involved a self-optimizing test. The idea was that, once we reached a certain confidence level, we would slowly grow the number of targets that were fed the most successful variant.
We had a minor (on the order off off-by-one or switching a < and a <=) code error that grew the successful variant too quickly, at a point where the confidence level was effectively non-actionable.
As I recall, it took us about six months to notice, and none of our clients noticed.
MVT, and especially our implementation, is obviously much more complicated than straightforward A|B testing. Given the fact that no one was able to sniff out such an obvious error when their tests didn't improve conversion as much as expected has left me with the idea that, while testing is not snake oil, I have 99% confidence the population involved in split testing has only a superficial idea of what they're doing.
[1] I had previously implemented a very simple Apache plugin, mod_gating, that I should clean up and throw on github. Most of the work was in the lexer for the configuration file. :-)
[2] Much of the design of appropriate statistical models was done through consulting with statistics departments at a couple local top-ten universities. We figured advanced stats is like cryptography, if you're not an expert in the general field and you come up with a "proprietary" solution, you're probably screwing something up.
I do understand statistics behind A/B testing quite well but there are many subtleties (especially related to interaction effects) that I still need to understand and learn on the way.