A Better Approach to A/B Test Analysis
blog.sendwithus.com
blog.sendwithus.com
We're both here to answer any questions :)
w/r/t testing multiple hypotheses, I've long wondered why false discovery rate never gets mentioned in the A/B-testing context. Any thoughts?
Could you explain why you've chosen to take a binary approach to determining whether differences are meaningful or not?
To me it seems like a continuous approach would be both more useful and more realistic. Creating an artificial threshold for significance seems a bit silly (and it also makes the model harder for users to use, because different applications might need different significance levels to justify an action).
From my perspective, every data point contains information and if you wait for significance you're essentially ignoring early information.
Edit: Also, when switching costs are small, significance levels become mostly pointless and you just want to switch to the best A/B option immediately. As evidence swings the other way, you just switch back.
I'd like to see someone takle creating a method aimed at our situations where results steadily trickle in. There ought to be a way to come up with adaptive thresholds such that at any given time we can ask, "Do we have statistically significant results yet, or do we keep the test running?"
Our Z-Test Method in Confidence.js pinpoints when results become statistically significant by calculating the required sample size at any given time using the standard error, zScore, and margin of error. If we have more samples than required, we can stop our test with confidence that we have statistically significant results.
We're still working out a way to do this with the new method - we have a hunch it might have to do with identifying when the confidence level becomes stable (as it does with more samples over time).
Here's a link to a paper on a Bayesian approach to the multi-armed bandit problem: http://onlinelibrary.wiley.com/doi/10.1002/asmb.874/abstract...
If you look at the history, I believe that you'll find that Pearson originally came up with the g-test as an approximation to an exact test, and then found the chi-square as an easier to compute alternative. This mattered back in the days of pencil and paper, but there is no excuse today to use the worse technique.
I'm going to avoid long discussions about the advisability of taking multiple looks at results with a classical statistical test. But see my incomplete series at http://elem.com/~btilly/ab-testing-multiple-looks/index.html for some of the considerations.
I never got into Bayesian statistics in there. In general they depend on the existence of a prior distribution. Careful treatments will talk about this. Sloppy ones assume one, don't talk about the one that they assume, and then quote results without letting you know about this important assumption. As long as you accept that assumption, they work well. But sometimes can be confusing to explain. (Until people "get" it. Then it can become irritating getting them to STOP explaining it!)
If you want to discuss these issues more, my email is my name at gmail.com.
Relevant keywords here include 'trial sequential analysis', 'adaptive trials', & 'multi-armed bandits'.
http://www.bayesianwitch.com/blog/2014/bayesian_ab_test.html
It can be drastically improved in the case of emails by building a Bayesian model tuned to email itself. Most people just blindly apply a testing procedure to their situation, but if you tune the test to your situation you can make do with a LOT fewer samples.
http://visualrevenue.com/blog/2013/02/tech-bayesian-instant-...