How Not To Run An A/B Test
evanmiller.org
evanmiller.org
I addressed this in my 2008 tutorial on A/B testing at OSCON. What I did is ran Monte Carlo simulations of running an A/B test while continuously following the results with various sets of parameters, and running the test to different confidence levels. In that model I peeked at every single data point. You can find the results starting at http://elem.com/~btilly/effective-ab-testing/#slide59. (See http://meyerweb.com/eric/tools/s5/features.html#controlchart for the keyboard shortcuts to navigate the slides.)
My advice? Wait until you have at least a certain minimum sample size to decide. Only decide there with high certainty. And then the longer the experiment runs, the lower the confidence you should be willing to accept. This procedure will let you stop most tests relatively fast, but still avoids making significant mistakes.
After that it is all a question of how many trials you were prepared to run, and how aggressively you need the answer. I tend to push to be conservative because the product manager is guaranteed to push the other way. For instance if you've got less than 500 trials I like to push for 99.9% confidence because, "If the effect is really this strong, we'll get there pretty quickly." After that I ease off. I've never tried to sit down and calculate any kind of optimal way to do so.
If I was to try to formalize it, I'd be inclined to set things up so that for some underlying bias which is smaller than what we're hoping to measure (for instance a 2% bias), our odds of picking the wrong answer are the same each time. I don't have any analysis behind that idea, it is just something that sounds reasonable to me. If I was to head down that road I'd probably run some Monte Carlo simulations to show that I'd be highly likely to wind up with test result answers in some vaguely reasonable time frame with whatever volume the website had. But again I haven't tried to do that analysis.
Just wanted to let you know that that slideshow changed my life. It made the company I'm with (FreshBooks) truly shine while doing split tests, which made me look good. Anyways, thanks dude.
I'm less mathematically sophisticated than the author, and would choose a simpler approach: ignore weak results. If one determines that there is a 95% chance that 51% of people prefer Logo A, either stick with what what you have, go with the one you like, or keep searching for a better logo. If you can't see the effect in the raw data without rigorous mathematical analysis, it's probably not a change worth spending much time on.
Instead of adjusting your significance test for each 'peek', simply ignore anything less than 99.9% 'significant'. And while you are at it, ignore anything that's less than a 10% improvement, on the assumption that structural errors in your testing are likely to overwhelm any effects smaller than this. Drug trials and the front page of Google aside, if the effect is so small that it flips into and out of 'significance' each time you peek, it's probably not the answer you want.
http://news.ycombinator.com/item?id=1203295
Or did I miss your posting?
The only danger is having a 'hidden variable' influence your results and averaging over the longer term masks that influence. For example, if you are not geo-targeting your content, you could conclude after a long run of testing that a certain page performs better than another, only to throw away the averaged out effect of having the different pages up on different times of the day, one of them performing significantly better for one audience and vv.
So you should keep all your data in order to figure out if such masking is happening and giving you results that are good but that could be even better.
The constraint here is not the math or technology rather it is users' needs. They want data, reporting and significance calculation to be done in real time. And even though we have a test duration calculator, I haven't seen any user actually making use of it. Plus many users will not even wait for statistical significance to be achieved.
Though, in VWO, we will love to wait calculating significance until end of experiment. I'm sure the users won't like it at all.
In my opinion, peak early and often and when your gut tells you something is true, it probablly is.
It is actually a mathematical fact if at any point in your A/B tests A is bigger than B, based on that data there is at least a 50% prob that asymptotically A is bigger than B.
"...when your gut tells you something is true, it probablly is."
If you run a data-driven business with a philosophy like that, you've rewound management science to about 1700 AD. Human "guts" aren't evolved for evaluating UX effectiveness from sparse data.
If you're going to take several peeks as you run your trial and you want to be particularly rigorous, consider alpha spending functions. In medicine, alpha-spending functions are often used to take early looks trial results. 'Alpha' is what you use to determine which P-values you will consider significant. To oversimplify a bit, early peeks (before you've got your full sample size) have very extreme alphas. If your trial ultimately uses an alpha of 0.05, a prespecified early look may use an alpha of 0.001. (There are ways of calculating a meaningful alpha values; these are just examples drawn from a hat.)
By setting useful alphas and betas, you can benefit from true, potent treatment effects (if present) earlier than you might otherwise, without too much risk of identifying spurious associations.
Now, in the world of the web where measurement has an upfront cost but 0 incremental cost, why not move to p < 0.001 or p < 0.0001? Sure, you need to increase the magnitude of data you're gathering by 2 or 3 but that's so much easier than delving into the epistemological complexities of p < 0.05
Don't stop before the test is complete, just because you've gotten an answer.
I generally leave my A/B tests up well after I've gotten a significance report, mostly because I'm lazy but also because I know that given enough time and enough entries, the significance reports can change.
Especially in the multi-variate tests that Evan wrote about, just because you get one result as significant doesn't preclude other possibilities from also being significant.