A/A testing
jvns.ca
jvns.ca
The "A/A" method described is not a terribly robust way to estimate variance, but the basic idea of using subsamples to estimate variance is what bootstrapping does more systematically.
There is even some recent work on a "Big Data" (distributed) version of the bootstrap from Michael Jordan's group [1]. It's also pretty easy to implement and can be really useful in practice.
[1] http://onlinelibrary.wiley.com/doi/10.1111/rssb.12050/full
More interestingly, things like revenue/visitor have a known probability distribution. It's not normal, but it is known. You can use a LOT fewer samples if you use a parametric test (either Bayesian or SPRT) based on the correct distribution.
If you use bootstrapping instead, you'll a) give up all your finite-sample guarantees and b) wind up using a LOT more samples than you need.
This actually depends. Based on my experience at a very early stage startup, it was definitely not the case for some datasets (I even tried fiddling with various well-known distribution's parameters).
If I recall correctly, I believe that the bootstrap had some asymptotic guarantee on the rate of convergence (although my memory is hazy on this)?
EDIT: never mind, it is asymptotic, hence not finite-sample necessarily.
If you compare the 2 batches of revenue/session distributions using a monte-carlo simulation you can calculate the probability that one is significantly different than the other. This generalizes beyond the 2 sample t-test because those underlying distributions are non-normal.
Please let me know if I'm thinking of this correctly (or not)
Thus, you can do all your normal Stats 101 tests on it.
If you compare the 2 batches of revenue/session distributions using a monte-carlo simulation you can calculate the probability that one is significantly different than the other.
A frequentist test (which includes most bootstrap methods) can never tell you this. Frequentist statistics doesn't even acknowledge this as a legitimate question to ask.
Now I agree, if you can use the exact distribution of revenues directly in the test, you can get answers even before you have enough samples for the CLT to apply. But if you use a nonparametric method like bootstrap, you'll need to use up a lot of samples unnecessarily.
I'm really nerd sniped here. Is there any branch of statistics that focuses on human understanding? For example, there's all kinds of blogs and stories out there about how doctors routinely make wrong choices because they don't understand statistics well enough. Is there any serious body of knowledge that explores ways of getting these doctors to make these mistakes less frequently, without having to send them to sites with titles like "An Intuitive Explanation of Eliezer Yudkowsky’s Intuitive Explanation of Bayes’ Theorem"?
What do you then? Well, the natural answer is to try the resample again. Do this n times, get an average for the variability, and that is precisely bootstrap.
At my previous employer, we open sourced a command line utility that we used to validate our statistical models if anyone's interested: https://github.com/monetate/monte-carlo-simulator
Offline resampling methods, like bootstrapping, are better if you're looking to robustly estimate the variance of the experiment.
You should also run an ongoing A/A test across your site or app to have confidence your bucketing, data pipeline, stats tests and effect on metrics is working as expected over time.
What is especially powerful about bootstrapping is that it does n't make any simplifying assumptions about the underlying distribution, unlike other methods to obtain confidence intervals.
What you really want are confidence intervals which show what would be a significant change. You can calculate that from your A-data and from your B-data. If they overlap you probably aren't quite there yet.
Comparing A/A vs. B or A/A/...A/A vs. B/B/...B/B is a poor man's approach to visualize the distribution of the mean values.
Things get further complicated when doing a lot of tests. If you do hundreds of A/B-Tests and a handful show a weakly significant result that may actually be a statistical fluke. The likelyhood that a wrong seemingly significant result is present when doing hundreds of tests can actually be pretty high. You should rerun these tests with fresh data and check for consistency, which in itself is some kind of A/A/B/B-test.
I always thought that statistical significance isn't something that should be tried to achieve, but merely a performance indicator how good the experiment was. Isn't it odd to try to "achieve significance over time"?
Shouldn't it be: "Your experiment requires 5,000 visitors and after that we'll check if the result was significant enough to not be merely due to random chance"?
Could someone with more statistical understanding elaborate this a bit?
That's basically what is happening with the tool, I think. It is asking for how many users per day you get in order to approximate the sample size for x days, then it's asking how much power you want. Power is the likelihood of detecting a difference if there is one. It also asks what confidence level you want. All of those together give you an approximate answer to the amount of time, assuming # of users/day is roughly constant.
You have to know all 4 before you do a test. A test is designed specifically to detect a certain difference. You cannot launch a test without knowing that as part of your hypothesis.
These are from the team that built amazon's weblab. The foundation of large scale web experimentation.
To be working in this field and not be familiar with this work, eg. the concept of A/A testing, is like deciding to build jet engines without having heard the idea of a bypass ratio.
A control group split into two is a good compromise, and intuitive to reason about, like the author points.
A quick and dirty way to avoid having to do much of any stats. Interesting.