The arithmetic is simple and cheap. Understanding basic intro stats principles, priceless.
The arithmetic is simple and cheap. Understanding basic intro stats principles, priceless.
Apparently, if you do the observing the right way, that is a sound way to do that. https://en.wikipedia.org/wiki/E-values:
“We say that testing based on e-values remains safe (Type-I valid) under optional continuation.”
[1] https://arxiv.org/abs/2210.0194
[2] https://www.evanmiller.org/sequential-ab-testing.html
[3] https://github.com/assuncaolfi/savvi/
If you are doing a lot of significance tests you need to adjust the p-level to divide by the number of implicit comparisons, so e.g. only accept p<0.001 if running ine test per day.
Alternatively just do thompson sampling until one variant dominates.
Thompson/multi-armed bandit optimizes for outcome over the duration of the test, by progressively altering the treatment %. The test runs longer, but yields better outcomes while doing it.
It's objectively a better way to optimize, unless there is time-based overhead to the existence of the A/B test itself. (E.g. maintaining two code paths.)
A key point here is that P-Values optimize for detection of effects if you do everything right, which is not common as you point out.
> Thompson/multi-armed bandit optimizes for outcome over the duration of the test.
Exactly.
In particular, if you aren't doing perfectly random sampling it is meaningless. If you are concerned about other types of error than sampling error it is meaningless.
A significant p-value is nowhere near proof of effect. All it does is suggestively wiggle its eyebrows in the direction of further research.
By "effect" I mean "observed effect"; i.e. how likely are those results, assuming the null hypothesis.
I find it hard to imagine obtaining much bias from a random hash seed in a large group of small-scale users, but I haven't looked at the problem closely.
I think this is not unrelated to the fact that if you wait long enough you can get a positive signal from a neutral intervention, so you can literally shuffle chairs on the Titanic and claim success. The incentives are against accuracy because nobody wants to be told that the feature they've just had the team building for 3 months had no effect whatsoever.
I’m not an expert but my understanding is that it’s doable if you’re calculating the correct MDE based on the observed sample size, though not ideal (because sometimes the observed sample is too small and there’s no way round that).
I suspect the problem comes when people don’t adjust the MDE properly for the smaller sample. Tools help but you’ve gotta know about them and use them ;)
Personally I’d prefer to avoid this and be a bit more strict due to something a PM once said: “If you torture the data long enough, it’ll show you what you want to see.”