Annoying A/B testing mistakes
posthog.com
posthog.com
> Bad Hypothesis: Changing the color of the "Proceed to checkout" button will increase purchases.
This is succinct, clear, and is very clear what the variable/measure will be.
> Good hypothesis: User research showed that users are unsure of how to proceed to the checkout page. Changing the button's color will lead to more users noticing it and thus more people will proceed to the checkout page. This will then lead to more purchases.
> User research showed that users are unsure of how to proceed to the checkout page.
Not a hypothesis, but a problem statement. Cut the fluff.
> Changing the button's color will lead to more users noticing it and thus more people will proceed to the checkout page.
This is now two hypotheses.
> This will then lead to more purchases.
Sorry I meant three hypotheses.
* Turns out, folks see the "buy". Many don't understand why they would want it. Some of those are converted after noticing and reading an explanatory blurb in the lower right. A more prominent "buy" button distracts from that, leading to more "no".
* For some reason, a flashing puke-green "buy" button is less noticable, as evidenced by users closing the window at a much higher rate.
Including untestable reasoning in a chain of hypothesises leads to false confirmation of your clever hunches.
We see a lot of ghosts in A/B testing because we are loosey goosey about our denominators. Mathematicians apparently hate it when we do that.
Would there be any benefit from knowing the notice rate though? After all, the intended outcome is increased sales by clicking.
Not sure if people do this, but with a mechanism of action in place you can state a prior belief and turn your AB testing results into actual posteriors instead of frequentist metrics like p-values which are kind of useless.
Also, frequentism / Bayesianism is orthogonal to causal / correlational interpretations.
Having a control doesn't mean you can't fall victim to this.
Observational ones are particularly prone because you can slice and dice the world into near-infinite observation combinations, but people often do that with AB tests too. Shotgun approach, test a bunch of approaches until something works, but if you'd run each of those tests for different significance levels, or for twice as long, or half as long, you could very well see the "working" one fail and a "failing" one work.
Requiring a user problem, proposed solution, and expected outcome for any test is also good discipline.
Maybe it's just getting into pedants with the word "hypothesis" and you would expect the other information elsewhere in the test plan?
if you have done that properly, why ab testing? if you did that improperly, why bother?
ab testing moves from an hypotesis, because ab testing is done to inform a bayesian analysis to identify causes.
if one knows already that the reason is 'button not visible enough' ab testing is almost pointless.
not entirely pointless, because you can still do ab testing to validate that the change is in the right direction, but investing developer time for production quality code and risking business to just validate something one already knows seems crazy compared to just ask a focus group.
when you are unsure about the answer, that's when investing in ab testing to discovery makes the most sense.
Except you can never be certain that the changes made were impactful in the direction you're hoping unless you measure it. Otherwise it's just wishful thinking.
but if you want to verify hipotesis and control for confounding factor, the ab test needs to be part of a baesyan analysis, if you're doing that, why also pay for the priori research?
by going down the path of user research > production quality release > validation of the hypotesis you are basically paying research twice and paying development once regardless of wether the testing is succesful or not.
it's more efficient to either use bayesian hypotesis + ab testing for research (so pay development once per hypotesis, collect evidence and steer into the direction the evidence points to) or use user research over a set of POCs (pay research once per hypotesis, develop in the direction that research points to)
if your research need validation, you paid for a research you might not need. if you start research knowing the priory (the user doens't see the button) you're not actually doing research, you're just gold plating a hunch, then why pay for research, just skip to the testing phase. if you want to research from the users, you do ab testing, but again, not against a hunch, but against a set of hypotesis, so you can eliminate confounding factors and narrow down the confidence interval.
As usual here I'm channeling David Deutsch's language and ideas on this, I think mostly from The Beginning of Infinity, which he delightfully and memorably explains using a different context here: https://vid.puffyan.us/watch?v=folTvNDL08A (the yt link if you're impatient: https://youtu.be/watch?v=folTvNDL08A - the part I'm talking about starts at about 9:36, but it's a very tight talk and you should start from the beginning).
Incidentally, one of these TED talks of Deutsch - not sure if this or the earlier one - TED-head Chris Anderson said was his all-time favourite.
plagiarist:
> That doesn't test noticing the button, that tests clicking the button. If the color changes it is possible that fewer people notice it but are more likely to click in a way that increases total traffic.
"Critical rationalists" would first of all say: it does test noticing the button, but tests are a shot at refuting the theory, here by showing no effect. But also, and less commonly understood: even if there is no change in your A/B - an apparently successful refutation of the "people will click more because they'll notice the colour" theory - experimental tests are also fallible, just as everything else.
Even though I agree, I'm not sure that's 100% epidemiology's fault by any means: it's just a very difficult subject, at least without measurement technology, computational power, and probably (machine or human) learning and theory-building that even now we don't have. But, there must be opportunities here for people making better theories.
This is called optional stopping and it is not possible using a classic t-test, since Type I and II errors are only valid at a determined sample size. However, other tests make it possible: see safe anytime-valid statistics [1, 2] or, simply, bayesian testing [3, 4].
[1] https://arxiv.org/abs/2210.01948
[2] https://arxiv.org/abs/2011.03567
[3] https://pubmed.ncbi.nlm.nih.gov/24659049/
[4] http://doingbayesiandataanalysis.blogspot.com/2013/11/option...
Anytime valid inference helps with this situation, but it doesn’t solve it. If you’re trying to detect a small effect, it would be nicer to figure out you need a million samples up front versus learning that because your test with 1,000 samples a day took three years.
Still, anytime is way better than fixed IMO. Fixed almost never really exists. Every A/B testing platform I’ve seen allows peeking.
I work with the author of the second paper you listed. The math looks advanced, but it’s very easy to implement.
Realize though that is a luxury, but I also see this trend in blue chip companies
As a data scientist at a "blue chip company", my team owns experimentation, but that doesn't mean we run all the experiments. Our role is to create guidelines, processes, and tooling so that engineers can run their own experiments independently most of the time. Part of that is also helping engineers recognize when they're dealing with a difficult/complex/unusual case where they should bring DS in for more bespoke hands-on support. We probably only look at <10% of experiments (either in the setup or results phase or both), because engineers/PMs are able to set up, run, and draw conclusions from most of the experiments without needing us.
I don't think A/B testing is a good idea at all for the long term.
Seems like a recipe for having your software slowly evolved into a giant heap of dark patterns. When a metric becomes a target, it ceases to be a good metric.
So, the button, (for user's most common resolution) had the button just below the fold.
This was an accident though some of our users called us out for it -- suggesting we'd removed the free plan altogether.
So, we a/b tested moving the button to the top.
It would REALLY hurt the bottom line and explained some growth we'd experienced. To remove the "dark pattern" would mean laying off some people.
I think you can guess which one was chosen and still implemented.
> more people signed up using Google and Github, overall sign-ups didn't increase, and nor did activation
Less friction on login for the user, 0 gains in conversions, they shipped it anyway. That's not a dark pattern.
If you're intentionally trying to make dark patterns it will help with that too I guess; the same way a hammer can build a house, or tear it down, depending on use.
Fundamentally, if you pick some number of metrics, you're always leaving some number of possible metrics "dark", right? Is there some objective method of deciding which metrics should be chosen, and which shouldn't?
Rolled out some tests to streamline cancelling subscriptions in response to user feedback, with Marketing's begrudging approval.
Short term, predictably, we saw an increase in cancellations, then a decrease and eventual levelling out. Long term we continued to see an increase in subscriptions after rollout, and focused on more important questions like "how do we provide a good product that a user doesn't want to cancel?"
Just don't test for dark patterns?
Because you're not testing for patterns, what you test is some measurable metric(s) you want to maximise(or minimise), right? So how can you determine which metrics lead to dark patterns, without just using them and seeing if dark pattern emerge? And how do you spot these dark patterns if by their very nature they're undetectable by the metrics you chose to test first?
Thanks.
Let's rephrase: if we are not to test for changes in user behaviour that give positive signal to progressive innovation, then what should we do? And how should we avoid the loudest voices in a room full of whiteboards creating a product that biases towards the needs of tech company and startup employees?
Should the passengers sign an agreement before being "experimented on", and having them be split in two groups, where one stays in a straight line and one in zig-zag?
Most of the a/b tests they reviewed (note the survivorship bias here, they were reviewed because they were surprising results) were incorrectly implemented and had to be redone. Most companies I worked at before or since did NOT have a team like this, and blindly trusted the results without hunting for biases, incorrect implementations, bugs, or other issues.
I can see extreme loads being valuable for an A/B test of a pipeline change or something that needs that load... but for the kinds of A/B testing UX and marketing does, leveraging statistical significance seems to be a smart move. There is a point where a large sample is trivially more accurate than a small sample.
That was a big one for awhile, and it would skew results.
Hmmm, another common one was doing geographic experiments when part of the experiment couldn't be geofenced for technological reasons. Or forgetting that a user could leave a geofence and removing access the feature after they'd already been given access to it.
Almost all cases boiled down to showing the user one thing while thinking we were showing them something else.
Number 2 (1 in the article) was solved by the platform. We had two activation points for UI experiments. The first was getting the users assignment (which could be cached for offline usage). At that point they became part of the test, but there was a secondary one that happened when the component under test became visible (whether it was a page view or a button). If you turned on this feature for the test, you could analyze it using the first or secondary points.
One issue we saw with that (which is potentially specific to this implementation), was people forgetting to fire the secondary for the control. That was pretty common but you usually figured that out within a few hours when you got an alert that your distribution looked biased (if you specify a 10:20 split, you should get a 10:20 ratio of activity).
Our approach to fixing these problems starts with having a golden path for running an experiment which essentially fits the OP. It's still going to take some work to educate everyone but the whole "golden path" culture makes it easier.
For giggles, we ran an analysis on those experiments: no difference between a & b.
That's usually the best result you can get, honestly. It means you get to make a decision of whether to go with a or b. You can pick the one you like better.
This is the biggest lie in experimentation. Of course you expect something. Why are you running this test over all other tests?
What I'm challenging is that if a team has spent three months building a feature, you a/b test it and find no effect, that is not a good outcome. Having a tie where you get to choose anything is worse than having a winner that forces your hand. At least you have the option to improve your product.
That's a great outcome. At one company we spent a few months building a feature only for it to fail the test, now that was a bad outcome. The feature's code was so good, we ended up refactoring it to look like the old feature and switching to that. So there was a silver lining, I guess.
The key takeaway was to never a/b test a feature that big again. Instead we would spend a few weeks to build something that didn't need to scale nor feature complete. (IOW, an MVP/POC shitty code).
If it had come out that there was no difference, we would have gone with the new version code because it was so well built -- alternatively, if the code was shit, we probably would have thrown it out. That's why its the best result. You can write shitty POC code and toss it out -- or keep it if you really want.
Because it has the best chance to prove/disprove your hypothesis. That's it. Even if it doesn't, all that means is that the metrics you're measuring are not connected to what you're doing. There is more to learn and explore.
So, you can hope that it will prove or disprove your hypothesis, but there is no rational reason to expect it to go either way.
> there is no rational reason to expect it to go either way.
Flipping a coin has the same probability of heads during a new moon as during a full moon. I’m going to jump ahead and expect that you agree with that statement.
If I phrase that as a hypothesis and do an experiment, suddenly there’s no rational reason to expect it to go either way? Of course there is. The universe didn’t come into being when I started my experiment.
Null hypothesis testing is a mental hack. A very effective one, but a hack. There is no null. Even assuming zero knowledge isn’t the most rational thing. But, the hack is that history has shown that when people try to act like they know nothing, they end up with better results. People are so overconfident that pretending they knew nothing improved things! This doesn’t mean it’s the truth, or even the best option.
You are trying to apply science to commercial applications. It works, but you cannot twist it to your will or it stops working and serves no purpose other than a voodoo dance.
> Flipping a coin has the same probability of heads during a new moon as during a full moon. I’m going to jump ahead and expect that you agree with that statement.
As absurd as it sounds, it’s a valid experiment and I actually couldn’t guess if the extra light from a full moon would have a measurable affect on a coin flip. Theoretically it would, as light does impart a force… but whether or not we could realistically measure it would be interesting.
Yes, I’m playing devils advocate, but “if the button is blue, more people will convert” is just as absurd a hypothesis, yet it produced results.
> You are trying to apply science to commercial applications. It works, but you cannot twist it to your will or it stops working and serves no purpose other than a voodoo dance.
With this paragraph, you’ve actually built most of the bridge between my viewpoint and yours. I think the common scientific method works in software sometimes. When it does, there are simple changes to make it so that it will give better results. But most of the time, people are in the will-twisting voodoo dance.
People also bend their problems so hard to fit science that it’s just shocking to me. In no other context do I experience a rational, analytical adult arguing that they’re unsure if a full moon will measurably affect a count flip. If someone in a crystal shop said such a thing, they’d call it woo.
Imagine a website is testing a redesign, and they want to decide if people like it by measuring how long they spend on the site to see if it's more "engaging". But the new site makes information harder to find, so they spend more time on the site browsing and trying to find what they're looking for.
Management goes, "Oh, users are delighted with the new site! Look how much time they spend on it!" not realizing how frustrated the users are.
EDIT: For some reason it didn't compute with me that you already referred to the same example. I've seen that exact scenario play out in real life, though.
"People spent more time scrolling the feed, people must enjoy it!"
No, the feed takes up more space, so now I can only fit 1 or 2 items on my screen at once, rather than 10, so I have to scroll more to see more content.
If only they were trying to maximise enjoyment and not addictiveness. They don't care at all about enjoyment, just like Facebook doesn't care about genuine connection to family and friends, or twitter to useful and constructive discussion that leads to positive social change.
I’m sure the first few of those got a lot of clicks, but it prompted me to ignore absolutely everything that comes from LinkedIn except for actual connection requests from people I know. Lots of clicks but also lots of pissed off people. I guess the latter is harder ti measure.
Edit - and this: http://www.stat.columbia.edu/~gelman/research/unpublished/p_...
Simpson’s paradox is about spurious correlations between variables - conversion analysis is pure Bayesian probability.
It shouldn’t be possible to have a group as a whole increase its probability to convert, while having every subgroup decrease its probability to convert - the aggregate has to be an average of the subgroup changes.
Consider the case where iOS users are more likely to convert than Android users, but you currently have very few iOS users. You then A/B test a new design that imitates iOS, but has awful copy. Both iOS and Android users are less likely to convert, but it attracts more iOS users.
The group as a whole has higher conversion because of the demographic shift, but every subgroup has less.
On that topic – what do you do when you observe that in your test results? What's the right way to interpret the data?
There are two issues at play here -- one is that the sample sizes for the segments may not be high enough, the other is that the more segments you look at , the greater the probability for finding a false positive.
At first, it would pick say a 50/50 split. Then as data rolls in that shows group A is more likely to convert, shift more users over to group A. Keep a few users on B to keep gathering data. Eventually, when enough data has come in, it might turn out that flow A doesn't work at all for users in France - so the ideal would be for most users in France to end up in group B, whereas the rest of the world is in group A.
I want the framework to do all this behind the scenes - and preferably with statistical rigorousness. And then to tell me which groups have diminished to near zero (allowing me to remove the associated code).
For example, you run two campaigns:
1. Get my widgets, one year only 19.99!
2. Get my widgets, first year only 19.99!
The first one may win, but they all cancel at the second year because they thought it was only for one year. They all leave reviews complaining that you scammed them.
So, I would venture that this idea is a bad one, but sounds good on paper.
[1]: https://medium.com/geekculture/the-novelty-effect-an-importa...
PS. A/B tests don't just provide you with evidence that one solution might be better than the other, they also provide some protection in that a number of participants will get the status-quo.
It's a great idea, it's just vulnerable to non-stationary effects (novelty effect, seasonality, etc). But it's actually no worse than fixed time horizon testing for your example if you run the test less than a year. You A/B test that copy for a month, push everyone to A, and you're still not going to realize it's actually worse.
My experience is that there's a good reason why this hasn't taken off: the returns for this degree of optimization are far lower than you think.
I once worked with a very eager, but junior DS who thought that we should build out a massive internal framework for doing this. He didn't quite understand the math behind it, so I build him a demo to understand the basics. What we realized in running the demo under various conditions is that the total return on adding this complexity to optimization was negligible at the scale we were operating at and required much more complexity than our current set up.
This pattern repeats in a lot of DS related optimization in my experience. The difference between a close guess and perfectly optimal is often surprisingly little. Many DS teams perform optimizations on business processes that yield a lower improvement in revenue than the salary of the DS that built it.
Is the idea that you wa t to optimize the conversion and then you would remove the experiment code with the winning variant ?.
Or would you prefer to keep the code in and have it continuously optimize variants ?.
I'd let the experiment framework decide (ie. optimize) who gets shown what.
Over time, the maintenance burden of tens of experiments (and every possible user being in any combination of experiments) would exceed the benefits, so then I'd want to end some experiments, keeping just whatever variant performs best. And I'd be making new experiments with new ideas.
[1]: https://github.com/lightswitch05/hosts/blob/master/docs/list...
As a company grows there will be multiple experiments running in parallel executed by different teams. The underlying assumption is that they are independent, but it is not necessarily true or at least not entirely correct. For example a graphics change on the main page together with a change in the login logic.
Obviously this can be solved by communication, for example documenting running experiments, but like many other aspects in AB testing there is a lot of guesswork and gut feeling involved.
It's important to not only A/B test minor changes, but occasionally throw in some major changes to see if it moves the same metric, possibly even more than your existing success.
"posthog.getFeatureFlag('experiment-key')"
doesn't look like it's actually performing a mutation.
If you're a medium-small business I see why you'd be tempted to roll your own. Trustworthy options under $15k/year are not apparent.
Isn’t the biggest problem with A/B testing that very few web sites even have enough traffic to properly measure statistical differences.
Essentially making A/B testing for 99.9% of websites useless.
1) https://experimentguide.com/ 2) https://bit.ly/ABTestingIntuitionBusters
> Good hypothesis: User research showed that users are unsure of how to proceed to the checkout page. Changing the button's color will lead to more users noticing it (…)
Mind that you have to prove first that this preposition is actually true. Your user research is probably exploratory, qualitative data based on a small sample. At this point, it's rather an assumption. You have to transform and test this (by quantitative means) for validity and significance. Only then you can proceed to the button-hypothesis. Otherwise, you are still testing multiple things at once, based on an unclear hypothesis, while merely assuming that part of this hypothesis is actually valid.
However, you should not dismiss qualitative results out of hand.
If you do usability testing of the checkout flow with five participants and three actually verbalize the hypothesis during checkout (“hm, I’m not sure how to get to the next step here”, “I don’t see the button to continue”, after searching for 30s: “ah, there it is!” – after all of which a good moderater would also ask follow up questions to better understand why they think it was hard for them to find the way to the next step and what their expectations were) then that‘s plenty of evidence for the first part of the hypothesis, allowing you to move on to testing the second part. It would be madness to quantitatively verify the first part. A total waste of resources.
To be honest: with (hypothetical) evidence as clear as that from user research I would probably skip the A/B testing and go straight to implementing a solution if the problem is obvious enough and there are best practice examples. Only if designers are unsure about whether their proposed solution to the problem actually works would I consider testing that.
Also: quantitative studies are not the savior you want them to be. Especially if it’s about details in the perception of users … and that’s coming from me, a user researcher who loves to do quantitative product evaluation and isn’t even all that firm in all qualitative methods.
Mind that 3/5 doesn't meet the criteria of a binary test. In statistical terms, you do know nothing, this is still random. Moreover, even if metrics are suggesting that some users are spending considerable time, you still don't know why: it's still an assumption based on a negligible sample. So, the first question should be really, how do I operationalize the variable "user is disoriented", and, what does this exactly mean. (Otherwise, you're in for spurious correlation of all sorts. I.e. you still don't know why some users display disorientations and others not. Instead of addressing the underlying issue, you rather fix this by an obtrusive button design, which may have negative impact on the other group.)
Everything you say is completely correct. But useful? Or worthwhile? Or even efficient?
The goal is not to find out why some users are disoriented and some are not. Well, I guess indirectly it is. But getting there with rigor is a nightmare and to my mind not worthwhile in most cases. The hypothesis developed from the usability test would be “some users are disoriented during checkout”. That to me would be enough evidence to actually tackle that problem of disorientation, especially since to me 3/5 would indicate a relatively strong signal (not in terms of telling me the percentage of users affected by this problem, just that it’s likely the problem affects more than just a couple people).
The more mysterious question to me would actually be whether that disorientation also leads to people not ordering. Which is a plausible assumption – but not trivially answerable for sure. (Usability testing can provide some hints toward answering that question – but task based usability testing is always a bit artificial in its setup.)
Operationalizing “user is disoriented” is a nightmare and not something I would recommend at all (at least not as the first step) if you are reasonably sure that disorientation is a problem (because some users mention it during usability testing) and you can offer plausible solutions (a new design based on best practices and what users told you they think makes them feel disoriented).
Operationalizing something like disorientation is much more fraught with danger (and just operationalizing it in the completely wrong way without even knowing) than identifying a problem and based on reasonableness arguments implementing a potential solution and seeing whether the desired metric improves.
I agree that it would be an awesome research project to actually operationalize disorientation. But worthwhile when supporting actual product teams? Doubtful …
This is actually the crucial question. The disorientation is an indication for a fundamental mismatch between the internal model established by the user and the presentation. It may be an issue of the flow of information (users are hesitant, because they realize at this point that this is not about what they thought it may be) or on the usability/design side of things (user has established an operational model, but this is not how it operates). Either way, there's a considerable dissonance, which will probably hurt the product: your presentation does not work for the user, you're not communicating on the same level, and this will be probably perceived as an issue of quality or even fitness for the purpose, maybe even intrigue. (Shouting at the user may provide a superficial fix, but will not address the potential damage.) – Which leads to the question: what is the actual problem and what caused it in the first place? (I'd argue, any serious attempt to operationalize this variable will inevitable lead you towards this much more serious issue. Operationalization is difficult for a reason. If you want tho have a controlled experiment, you must control all your variables – and the attempt to do so may hint you at deeper issues.)
BTW, there's also a potential danger in just taking the articulations of user dislike at face value: a classic trope in TV media research was audiences critizising the outfit of the presenter, while the real issue was a dissonance/mismatch in audio and visual presentation. Not that users could pinpoint this, hence, they would rather blame how the anchor was dressed…
It is alright to decide that in certain cases you can act with imperfect information.
But to be clear, I actually think there may be situations where pouring a lot of effort into really understanding confusion is confusion. It‘s just very context dependent. (And I think you consistently underrate that progress you can make in understanding confusion or any other thing impacting conversion and use by using qualitative methods.)
The alternative is just running the experiment, which would take 10 minutes to set up, and see the results. The A/B test will help measure the qualitative finding in quantitative terms. It's not perfect but it is practical.
Wouldn't it be better to have an A/B testing system that just counts how many users have been in each assignment group and end when you have the required statistical power?
Time just seems like a stand in for "that should be enough", when in reality you might have a change in how many users get exposed that differs from your expectations.
However, the deceptively similar scheme of running it until the results are statistical significant is not okay!
Need I say more? Or just keep tweaking your website until it becomes a mindless, grey, sludge.
something new performs better cause its new
if you have a platform with lots pf returning users this one will hit you again and again.
so even if you have a winner after the test and make the change permanent, revisit it 2 months later and see if you are now really better of.
all changes of a/b tests in sum has a high chance to just get an average platform in the sum of all changes.
I hesitate to write this (because I don't want to be negative) but I get a sense that most software "engineers" have a very narrow view of the industry at large. Or this forum leans a particular way.
I haven't A/B tested in my last three roles. Two of them were defense jobs, my current job deals with the Linux kernel.
I don't work on the kernel, but one of the most professionally useful talks about the Linux kernel was an engineer talking about how to use statistical tests on perf related changes with small effects[1]. It's not an _online_ A/B technique but sometimes you pay attention to how other fields approach things in order to learn how to improve your own field.
Doesn't the online nature of an A/B automatically account for this?