A/A Split Testing
getelastic.com
getelastic.com
I don't understand why they find this surprising. Of course there's going to be some variation in the conversion rates. This is the reason why GWO reports statistical significance.
Is the method by which you're distributing sessions into your tests somehow biased?
I'm going to call A' the "new" (but identical) test, and A the original test.
If you've been running A for a long time, and now add A', what is the chance that the visitor populations between A and A' will be different enough to drive statisically significant differences that are population-related rather than test-content-related?
Put slightly differently: If your returning session conversion rate is higher than your first session conversion rate, you will need to take some pains to ensure that each of the tests is getting a fair shake at the traffic. In many cases, that means ending test A, and creating a new test a and a' such that neither a nor a' has an advantage. It's easy to ASSume that there is no meaningful test bias, while the reality is that it's quite easy to have test bias creep in.
Something I learned from years as a QA lead: always ask that every test you do has a failure and a success case. Surprisingly often, the tests that show your program works fine aren't actually doing their job properly, so they also need to be checked.
In other words, testing your implementation of A/B testing is a Good Thing. Even if you don't understand that statistics behind it, at least you can see the same variance occurs in A/A testing, which tells you to go read up some more on what you're doing wrong.
They hear about it, run off and implement the a/b versions of their pages and add some tracking. Proceeding then to draw incorrect conclusions and make bad decisions based on them. They don't go read up on it enough to understand the statistics behind it. I used my stats book from Uni (I knew it had a better use than a laptop stand, a java book now has that honour) but there are plenty of tutorials out there. For example: http://visualwebsiteoptimizer.com/split-testing-blog/what-yo...
It really is a toss up between statistics and medicine for which is the most misunderstood of the sciences.
That's why sites like Visual Website Optimizer are getting increasingly good. They use plain language to explain when to keep running the tests.
Again: It's unrealistic to expect lay users to learn statistics. They don't want to become statistical experts, they usually just want to increase sales. This is where software can help.
As an aside, wow another one of my favourite bloggers responding to a comment of mine on HN this week!
EDIT: I just noticed your username and I realise that you probably already understand this :p I still think it's a bit of an odd way of putting it though.
The question is whether it is statistically significant.