How we improved signups by 30% by doing nothing.
blog.historio.us
blog.historio.us
So, your 99.8% confidence isn't 99.8%.
There are a few ways to compensate for this. The easiest is to fix your sample size and not reach any conclusions until you have tested that many users.
You can determine your sample size by, say, figuring out the minimum number of people needed before there's a 95% chance you observe a 10% effect size.
This is a problem in clinical trials where ethical questions arise. If the test appears harmful, should we stop? If it appears beneficial, isn't it unethical to deprive the control group of the treatment?
Anyhow, if you want to observe the results continuously, you need to use a technique like alpha spending, sequential experimental design, or Bayesean experimental design.
TL;DR: If you're periodically looking at your A/B testing results and deciding whether to continue or stop, you're doing it wrong and your significance level is lower than you think it is.
Imagine that you wanted to see if A != B, with a 95% confidence. To oversimplify, this means that you're willing to accept that 1 in 20 times, you'll incorrectly reject the null and you will consider A different from B even though they are truly the same.
If you run 20 independent tests at once, each at a 95% confidence, then by chance you'd expect 1 to reject the null even if they're actually all null.
Now, if you repeatedly peek at the data as it accumulates for one test, you're doing something similar (not 100% the same because the tests aren't totally independent, but similar). To again oversimplify, you'd expect, by chance, to see "significant" results 1 in 20 of the times that you look at the data, even if there is no significant result. This is why you need to wait until you've collected all of the necessary data first, or implement a procedure to protect you from incorrectly seeing "significance" when it's not there. You might, for example, require a more stringent confidence level earlier on, and less stringent ones later.
"Statistical significance", not so much. A simple test program I wrote showed that with a limit of 500 coinflips, flipping a fair coin will produce, on at least one out of the steps, a dataset which rejects the null hypothesis with p < .05, fully 30% of the time.
http://lesswrong.com/lw/1gc/frequentist_statistics_are_frequ...
The meaning of an experimental result should not depend on the experimenter's private state of mind (such as the stopping rule they followed) if that state of mind does not affect the experimental apparatus. The likelihood ratio is objective, "statistical significance" is not. See link.
Well, keeping in mind that experimenters should just report likelihood ratios, I'd have estimated a pretty low prior probability that one of two identical pages produced more signups than the other.
But I suppose that after seeing a likelihood ratio of 1000/1 favoring the hypothesis "each of these pages has a separate fixed conversion rate" over "both of these pages have the same fixed conversion rate", using Laplace's Rule of Succession (i.e. a flat prior between 0 and 1 for the conversion rate) in all cases, I'd think that something was more likely than not going on.
For a new HIV test, I want to know the likelihood ratios, not other parameters which tend to be driven by the background HIV prevalence in the researcher's specific catchment area. My catchment area's HIV prevalence will differ, so other aspects of the test that are dominated by prevalence will not be the same in my population. However, the likelihood ratio will be.
The point that I didn't finish making before getting kicked out of the coffee shop was that, at some point, people will want to compute a posterior, so they'll need to figure out a prior. Your suggestion above is helpful.
Edit - Also, you're right in that I saw just the 1st paragraph when I initially replied.
As long as you decide on a minimum number of trials to make a decision, it's fine to look at the data as it is accumulating.
The exact numbers depend on how often you test the confidence interval, but a multiple of 5-10 is for early measures will usually give reasonable results.
E.g., let's say you ultimately want a 98% confidence interval (2% chance of error). A 10x improvement is a .2% change of error, or 99.8%. Depending on how often you are sampling, an early result with 99.8% confidence can be equally accurate as a later results with 98% confidence.
Not very many people have studied probability, whereas a lot (relatively) have heard the good effects of A/B testing.
I think A/B testing is a very good method, but you need a lot of data. I'd say, when you're just starting up, don't try A/B. You won't get the data you need.
It's very easy to be seduced by statistics. It doesn't matter if the stats are wrong.
EDIT: seduced, not deduced.
Which is more likely: (a) that this was the 1 in 1/(1-0.998) times that the correct method incorrectly rejected the null; (b) that there is actually a cryptic difference between A and B in the AB test; or (c) that the wrong method was used, or the right method was misused, or something along those lines?
I would say that those are arranged in increasing order of likelihood.
And that's why I liked this post -- it's highlighting the fact that it's easy to misread statistics.
This is what most people who talk about A/B testing don't mention, that you need more data than you think, in order to get good results.
This should not reduce your belief in confidence intervals; this is a great, motivating opportunity that should prompt you to learn how to use them correctly.
Thanks for the help!
This is actually terrible advice because continuing a test in which one set is significantly better than another has a cost. You are showing an inferior set to a segment of your users and that costs you money (or signups or whatever metric it is you're improving which, at the end of the day, presumably equates to money).
As an example, suppose you do a test and discover something that doubles your signup rate (and therefore monetization rate) and you've got a confidence level of 99.9%. It's true, there's a 1 in 1,000 chance your result is flawed and you'll end up with the wrong decision. But there's a 999 in 1,000 that you're showing a significantly inferior signup page to half of your customers, costing you about 25% of potential revenue. It doesn't even take someone who knows what EV stands for to realize his EV on ending the test is huge here.
You say, "How many people should I test before there's a 95% chance I observe a 10% effect size?"
That's your sample size. It's easy to compute up front.
But if I flip a coin 10 times and get the queens face (I'm in the UK) 8 of those times, that doesn't mean I'll keep getting head.
It's been quite a lot of years since I was involved in the betting world now, but at the time I was constantly amazed at the amount of people that thought the Martingale system was a winner.
I know I almost signed up but i'm still tossing it up between pinboard and historius