The Surprising Power of Online Experiments (2017)
hbr.org
hbr.org
Google has a pretty good FAQ on this:
https://support.google.com/analytics/answer/2847021?hl=en&re...
[0]: https://gtagency.github.io/2016/experimentation-with-no-ragr...
Definitely the way to go for A/B.
We've talked to Optimizely but their pricing was going to come in at the same ballpark as our AWS spend (into the six-figure range), which seems absurd. They charge based on monthly users, but a lot of our traffic consists of organic search bounces.
For now we just want to run ~5 experiments per month, want to record events server-side so we can be sure not to lose any, and are wary about implementing it ourselves since there are so many ways to screw it up without realizing it.
PS: it’s a product that I launched here on HN 9 years ago and wouldn’t have been possible without the awesome community.
1. If the effect size is incredibly small, which most minor UI changes will be, finding statistical tests to prove them is really difficult. If you’re looking for an incredibly small positive effect, even with hundreds of thousands of sample size, the probability of rejecting the null hypothesis while the effect is actually a negative influence is surprisingly high! Very easy to make mistakes.
2) short term gains on engagement may lead to long term disengagement.
3) business incentives for management are easily misaligned. I would imagine a dominant negative influence is managers exaggerating the statistical influence found in a test because that means they get to lead the change, on an otherwise vast tech ecosystem the performance of which probably won’t change all that much. Attribution is also hard (how sure are they on how much to attribute here?) so credit is difficult to allocate beyond initial value sizing.
I needed to find cites, and ironically I found the Parent article this morning. This one is way better about the lies we tell ourselves: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3204791
Once upon a time we published over 100 a/b tests here: https://www.goodui.org/evidence/ and clearly the relative effects vary (not all single changes have always a small effect).
More so, the effects of a/b tests can be further increased by grouping multiple higher confidence ideas together into a single variation.
2) Short term gains may (or may not) lead to long term disengagement. Measuring micro (shallow) and macro (deeper) metrics would be the right way to answer this.
I used to work with Ron Kohavi in his group.
Thanks for the heads-up.
As has been mentioned previously, our business depends on the trust users place in our services to provide reliable, high-quality information. The primary goal of our recommendation systems today is to create a trusted and positive experience for our users. Ensuring these recommendation systems less frequently provide fringe or low-quality disinformation content is a top priority for the company. The YouTube company-wide goal is framed not just as “Growth”, but as “Responsible Growth”.
[1] https://blog.google/documents/33/HowGoogleFightsDisinformati...
Someone could change the CSS of the ads, for example to flash periodically, and that would be the reason for the increased clicks.