A/B test improved your website's conversion rate? Not so fast
blog.alexandervolkmann.com
blog.alexandervolkmann.com
You can also run into this sort of problem with user learning effects, where initially a large change in the UI can give a large change in behavior due to novelty, but then it wears off over time. Running experiments longer helps a lot in both cases.
Rather, these are simulated data for a fictitious company. The author is demonstrating a scenario in which a purely frequentist approach to A/B testing can result in erroneous conclusions, whereas a Bayesian approach will avoid that error. The broad conclusions are (as noted explicitly at the end of the article):
- The data generating process should dictate the analysis technique(s)
- lagged response variables require special handling
- Stan propaganda ;) but also :(
It would be cool to understand what the weaknesses or risks of erroneous conclusion to the Bayseian approach in this or similar scenarios. In other words, is it truly a risk-free trade off to switch from a frequentist technique to a Bayesian technique, or are we simply swapping one set of risks for another?
tl;dr The author's point is not to make a general claim about the aggressiveness of CTAs.
While it is possible to make some progress on this issue with careful math, simply running the test longer is a far more effective and robust approach.
Also you and GP are calling the example fictitious, but seems to based on 'real traffic logs' via https://dl.acm.org/doi/10.1145/2623330.2623634
> "Let us consider the following fictitious example in which Larry the analyst of the internet company Nozama"
Nozama is Amazon backwards.
And to expand on this, the data generating process is not about a statistical distribution or any other theoretical construct. Only in the frequentist world do you start with assuming a generating process (for the null hypothesis, specifically).
The data generating process in this case are living, breathing humans doing things humans do.
The potential outcomes are fixed: if a person is assigned to one group the outcome is x1; if another, x2. No assumption is made about these potential outcomes. They are not considered random, unless the Population Average Treatment Effect is being estimated. And even in that case, no distribution is assumed. It certainly is not Gaussian for example.
Under random assignment, the observed treatment effect is unbiased for the Sample Average Treatment Effect. So again, the data generating process of interest to the analyst is random assignment.
And it's my fault for not thinking of that as a possibility. Colour me jaded after experiencing very many bad attempts at randomization that actually suffer from Simpson's paradox in various ways!
You need to have basic intuition for what might be happening, the math is just a formality and frankly, is unnecessary beyond a very, very simple calculation.
You have to actually 'think' about behaviour a bit if you want to get it right, that's the hard part.
If you have something reasonable, then the conversions/control numbers can be worked out into a probability of success very quickly, and even then, just looking at them will give you a good idea if it worked or not.
The maths is a shiny lure for technical people, it gets us all excited as though there is some kind of truth behind it.
Sites can only optimize for what they can see and we've made it so they can only see short-term engagement.
Another is all the annoying cookie popups as a result of GDPR.
If sites are having trouble converting me, perhaps it's not me that's the problem.
It's somewhat amusing that the overlap of garbage content farms and sites with annoying consent popups is almost perfect. I wonder if it could be used for search engine ranking.
Previous testing should give the company at least some baseline understanding of what is trivial and what isn't. The correct way to experiment is certainly not "let's experiment every idea!"
> however obviously beneficial
If you've been around long enough, you've almost certainly run into dozens of "obviously beneficial" changes that led to poorer performance.
Most of what you're describing is issues with poor prioritization, a lack of understanding about your audience, and a culture that has a difficult time making decisions.
Metrics are only useful if the organization is actually willing to learn lessons from them.
Any time someone wants to measure something, the top two questions should be "what are the lower and upper bounds this value has to exceed for us to do something different?"
Very often, it turns out these thresholds for change are so astronomical that nobody thinks we have even the slightest chance of exceeding them. That means the measurement is completely useless. Whatever result we plausibly get, it won't change anything.
You can save a lot of time this way!
The reason I wanted this type of test was because it was a waste of time testing shades of blue or two headlines that only differed by 2 words. The test variants were never radical enough to see any kind of significant uplift. Then after 5-10 tests the design starts to suffer by wandering down some weird path that nobody would consciously design from the outset. But the series of test "winners" made things go off in wild directions.
I still think there is some value in A/B testing (A/A/B only, if I'm honest). But in a small team, it's a waste of time.
If A' and B both statistically differ from A, then you have a problem because you're not testing what you think you are testing, regardless of what your naive A/B test's p-value would have indicated.
This helps to show the effect of a low sample from a non-uniform distribution.
A lot of people (me included) think they know statistics, but they don't.
The blogpost in OP also tries to explain the same thing - you shouldn't do statistics without understanding what it is your doing.
1) It's almost always a bad idea to decide a test based on one-week's worth of data, regardless of what statistical approach you take
2) There's not really any info on why Fisher's exact test is used. It seems like most A/B software has adopted Bayesian but the ones who haven't, I believe, choose the Student's test and require prior sample sizing
3) The conversion delay issue was not addressed in the measurement plan. There are clear ways to address this issue, both tactically as well as mathematically. From a tactical standpoint, most testing platforms, you'd be able to change the test allocation to 0%, which would allowed previously bucketed users to continue to be measured with subsequent visits while not allowing any new users in. You could also just run the test long enough to where the conversion lag no longer has a major impact on results (this may or may not be possible, depending on how long and fat the lag tail is).
This can be quantified by plotting the incremental conversions observed by day x. We migh see a big initial lift that degrades over time. If it eventually degrades to zero, there are no truly incremental conversions, just pull-forward. But if we end up pulling forward a meaningful number of purchases by a month or more, that can be valuable to the business!
I wouldn't immediately jump to a complicated mathematical model to handle this situation, I would consider the business implications first and foremost.
I also urge anyone considering Bayesian methods for A/B testing to read up on the likelihood principle vs the strong repeated sampling principle (I documented my thoughts here [0]). Bayesian methods always satisfy the likelihood principle; frequentist methods always satisfy repeated sampling. In many situations both methods satisfy both principles, and then the two approaches will give similar answers. But based on many years doing A/B testing, I wouldn't give up repeated sampling lightly. Bayesian and frequentist methods are not blindly interchangeable.
On the other hand, if repeated sampling is not important in your use case, then by all means prefer the Bayesian approach! I just want people to consider the trade offs.
This basic principle has far broader implications than website design and A/B testing. Managers at large corporations have learned to pull all sorts of levers to optimize the short term value of some metric (typically the one upon which their compensation depends) often in direct opposition to the long term interests of the corporation (and even the value of that metric beyond the next few quarters).
[0] https://github.com/CamDavidsonPilon/Probabilistic-Programmin...
They are based on his text book with the same title.
(Disclaimer: I'm the author of the blog post.)
Is it even fine to use the distribution assumptions in the later analysis?
Looks like these assumptions combined with a higher conversion rate on day 2 for control is the main reason for the surprising result (control is obviously spread out).
The (fictitious) signal they are discussing here is very strong. Scroll down to the figure labeled "posterior distribution of p" and you can see that the two distributions barely overlap.
To me it just looks like a whole new batch of assumptions. Might be fictitiously valid or not.
And then there's the big question: How much of business did you lose in the process of arriving at what seems like an optimal solution (which might just be a local peak, rather than a global optimum point)?
That said, what's the alternative? To optimize, or not, that is the question.
Just because one design converts more than the other doesn't mean it's the design with optimal UX. I've seen many tests where the designs included already had faulty UX. This is why it's better to have a trained UX designer on your team who can fix basic flaws and present the best version of various designs for testing.
On the other hand, if the variants mostly perform the same, why spend more time on it? Go focus elsewhere.
It certainly is a logical possibility that the next design you try will be much more impactful, but after trying several variants unsuccessfully your time is probably better spent elsewhere.
If he went on to post about his preference on the Internet with some made up examples I would be sure that he couldn't be trusted.
It may be valid. It's insufficient as it stands to make a determination.
TL;DR: it's noise.
Does this model require a proper prior on p?
Cool blog post.
- I don't think a proper prior is required.
- Thank you :)
(Disclaimer: I am the author of the blog post.)
I'm not sure I've ever encountered the second issue with screen width.