The multiple comparisons problem
chrisstucchio.com
chrisstucchio.com
They might be from the marketing department, they might be web-designers, they might even be growth hackers (eek), but most likely they are not statisticians.
Their motivations may not align with running tests correctly. E.g. they may want to strengthen the argument for their design over someone else's competing design, or they may be looking for number to serve up to their boss at their monthly review meeting. But performing a statistically-valid test is low on their list of priorities.
I love articles like this that go into the minutiae of how to perform AB tests correctly. But I think they only speak to a small portion of the people who are doing AB tests.
I actively need to prevent myself from going too deep into the maths and caveats of A/B testing with clients. Running tests with bad maths trumps running no tests at all in my experience
I think there is a lot of pseudo-scientific talk around A/B testing that leads people to either put too much credence in their results, or, paradoxically, avoid advanced techniques due to fear their results will be inaccurate.
I'm coming to the conclusion that one should treat A/B testing from more of an economic point-of-view. You have so many views this month -- how are you going to spend them on tests? I think one should look at A/B tests as a source of information, not of truth, and thus avoid chasing after statistical rigor that won't, in fact, be realised in most cases.
Update: I should add I think the issues are more complex than simply using, say, Thompson sampling.
I'm curious, did our conversations a few years ago help shape this opinion? My guess is that we've both moved somewhat towards a middle position.
It might be of interest to you that there is now some published work on bandit algorithms for when all you have is ranking, not absolute ordering. This is close to what we were talking about: http://arxiv.org/pdf/1312.3393.pdf
Take the case given in the article: 2 designs that appear equal, but segmenting reveals (incorrectly) that one performs better on Android and the other on iPhone. Statistically, it doesn't matter which one you pick, because you don't have a strong result. Anecdotally, you might go down the wrong path and waste some time doing something like showing design A to Android and B to iPhone. But the only thing lost is time -- probably a small amount of time compared to creating the different designs.
Best case, you realize there's something technically wrong with your changes in a certain segment because things are widely different. Worst case, you make some changes that don't matter. Overall, the risk / reward of segmenting seems worthwhile.
But from a business perspective, it's a huge distraction. Your marketers shouldn't be wasting time segmenting - there is no reason to believe it's making any money. And due to multiple comparisons, it's virtually guaranteed that your A/B tests will steer the marketers in the direction of more segmentation.
Automatically ignore all statistics on segments that are below a certain sample size. Say, 4k successful conversions. If there are no real differences, you'll still draw lots of wrong conclusions. But your odds of wrongly drawing significantly harmful decisions will be very low, and some of the conclusions you draw will be very good.
Your procedures may make statisticians curse and swear. But it provides a simple rule of thumb that is easy for non-statisticians to understand. And mitigates the really bad mistakes.
But yeah, that is a pretty good way for a marketer to avoid being too stupid. And I can't say I haven't done this a bunch of times, particularly when doing quick exploratory calculations.