A/B testing and the historic lift paradox
bytepawn.com
bytepawn.com
Suppose that most of your revenue comes from recurring subscriptions of varying prices. Then, to use an extreme example for clarity, suppose that last month's revenue per person for the control and treatment groups respectively was not $95 and $101 as in the article, but $10 and $100. I guess you just got really, really unlucky with your random selection, such that most of your high-paying users landed in the treatment group. Then you wait a month and measure again, and it turns out this month's revenue per person was exactly the same in each group, $10 and $100.
By this article's logic, you should ignore the past values and only look at this month's $10 and $100, and conclude that the treatment made revenue 900% higher. In reality, the results probably just mean that most people didn't change their subscription level from month to month. Because individual users' spending on the subscription is highly correlated from month to month, you should measure this month's spending compared to last month's baseline.
A few others have mentioned CUPED (Kohavi et al. 2013), a way of doing exactly what this blog post says you can’t. CUPED is a great starting point for anyone new to the subject of pre-experiment control variates.
The post is bold enough to describe any use of control variates is a “fallacy” but somehow doesn’t mention the most famous method for doing this (CUPED). The author is only interested in linking to their own prior blog posts, and does not engage with any reasonable (or attributed) argument for the method they’re dismissing.
The post’s main argument is that using CUPED-style control variates provides no benefit when N is sufficiently large, so we should just make N bigger — as though “sufficient N” grew on trees.
The next argument is less silly, but still wrong: They note observed differences in pre-treatment effects are “random fluctuations”, and claim they thus cannot be used for anything. Their argument ignores the reason pre-treatment measurements are subtracted in the first place: In CUPED, we acknowledge pre-treatment observations are undesirable noise, but presume this noise already contaminates our post-treatment measurements. We collect pre-treatment measurements exactly because we want them gone — the pre-treatment measurements tell us how much meaningless noise to subtract.
As a reductio ad absurdum, imagine an “A/B test” (an RCT) of a drug that ostensibly makes humans taller. Would it seem reasonable to you, as a participant, if the doctors never ask what your initial height is, and only measure you once the trial is done?
After a quick scan, it seems to me that CUPED is supposed to be run at a per-unit level (ie. per user normalization), to reduce variance, which seems to be a bit different than the fallacy I describe here (computing lifts from the T and C group's "before" and "after" overall mean separately and subtracting the lifts, which is what I observed and triggered me to write this post).
(I wrote the post.)
Insofar as I think your intuition is leading you somewhere, I think it's leading you towards a realization that a "diff in diff" approach rather than regression adjustment can increase variance in some settings. But regression adjustment is provably better in essentially all circumstances: the only settings in which it is ever worse than no adjustment are outlined clearly in https://projecteuclid.org/journals/annals-of-applied-statist...
But you need to look at the "before" to show that your control and treatment are similar enough to compare.
If you randomly assign people to the test- then control and treatment should be the same in the "before"
There even exist procedures to rerandomize if pre-test metrics are lopsided. You retain the benefits of randomization and it can help a little when your signal is the same magnitude as your sampling noise.
Presumably, a high standard deviation of multiple randomised pre-groupings could show that your A/B testing is likely to be problematic.
The historical difference in average is much larger than the after treatment. Can those groups even be compared after treatment?