I admit to attempting to be guilty of this in the past and being stopped by our analytics team (in the sense that they took the time to patiently explain to me why what I was doing was statistically unsound). It's not obvious, IMO.
I admit to attempting to be guilty of this in the past and being stopped by our analytics team (in the sense that they took the time to patiently explain to me why what I was doing was statistically unsound). It's not obvious, IMO.
If you create a culture where positive A/B tests are lauded (which is good!), then you create a lot of people who want A/B tests to finish in positive ways. For those people, it doesn't really matter if their A/B test actually improves things, only if it looks like it does. This isn't nefarious, this is just human nature, but it creates a lot of creativity and energy at finding ways to making winning A/B tests. We'd have people run a test, where their new variation got off to a bad start, then say "oh, it was a bug", then they'd fix some irrelevant thing and start over just to reset the counters. I was guilty of it sometimes. You get excited about your tests and want them to win. That is why it is critical either to have gatekeepers like an analytics team to keep you honest or have a really specific protocol on how your company runs tests and only consider results of tests that followed the protocol.
There is statistics for the purpose of uncovering Truth, and statistics for the purpose of making a business decision. The difference is that when we talk about Truth, a small error is still an error. When we make business decisions, it is fine to make a decision that is probably right, and we know isn't far wrong.
Here is a perfectly valid test procedure that illustrates the difference. Decide the most time you would be willing to spend to get a test result. Multiply that by current conversion rates to get N, the number of conversions that you expect to see by the end of the test.
Start running the test with two variations. Stop at any point if one variation is at least sqrt(N) conversions ahead of the other. Stop at N if there is no clear winner and go with whoever is ahead, even by a hair.
Here are features of this test procedure.
o You always make a decision.
o Running a test has a known fixed cost. You know how long it takes. And a bad idea will cost you no more than sqrt(N) conversions to test.
o The results are very simple and easy to understand.
o Your answers are usually right.
o Your bad decisions are not very bad. If the true conversion rate for one version is better by 1/sqrt(N), you've got a 95% chance of making the right choice. You will probably never make a mistake as big as 2/sqrt(N).
The result is a test procedure that is a horrible approach for doing science, but an excellent tool for improving a business. You'll never find it in a statistics class. And I'm sure it would horrify your analytics team.
So I think its wrong to say that you'd never find it in a stats class.
In bandit there is a clear explore/exploit trade off. There is no such trade off in the A/B formulation, although it does get used in scenarios that have such trade offs.
If I can pull the lever a finite and small number of times there is a strong incentive for using a bandit. In this case I don't want to pull the wrong lever as few times as possible. On the other hand if I am given an unlimited number of pulls, I can afford to pull the wrong one many times more (still finite) for the sake of 'knowledge' knowing well that I would have infinitely many opportunities to exploit that knowledge.
In other words, the opportunity cost of putting someone in the wrong group creates such a trade-off. You can pull the lever as many times as you want, but each one potentially costs you money. It's textbook bandits.
Differences creep in when there is ambiguity and judgement involved on what is that metric that the org wants to optimize. This is fairly common. Typically, in these situations its the PMs who make the final call. There the goal of the experiment protocol is glean as much knowledge as possible, and present it to the PM. The thinking there is -- if that comes at the cost of exposing some customers to bad choices, so be it.
If A is ahead of B by a hair and the my flawed protocol chooses B the cost to business might not be that high. But the same protocol might not be a good one if the cost of making that mistake is very high. The probabilities of the errors remain the same for the two scenarios, the expected costs are different.
Honestly, for most startups a simple multi-armed bandit approach is probably the way to go. Don't worry about statistical significance; just throw some "lite" reinforcement learning on top of your product's aesthetics and enjoy the incremental profit. (Caveat: do not apply MAB to major product changes.)
Doesn't this mean maintaining all the variants forever?
But yes, statistics is not obvious at all.