An A/B Testing Story
training.kalzumeus.com
training.kalzumeus.com
Spot on, and you should 100% support Patrick & Nick and buy their books / videos / etc.
But A/B testing is not as simple as hiring an accountant, running A/B testing as part of your corporate culture requires:
- Front end developer to install / manage testing engine.
- UX designer to measure / understand where holes may be in your flow.
- Competent web designer to execute instructions from UX designer.
- Statistically sound processes to know when to "call" a test (no Optimizely does NOT do this)
- SEO / marketing people to make sure that your A/B test didn't just break marketing flow
- etc
Yes, anybody could go run a headline test.
But running A/B testing with any sort of regularity, scale, and success is a complicated problem.
Assign visitors into buckets randomly. Track conversion counts only. Declare the winner of an A/B test whenever one gets 100 conversions ahead, or whichever is ahead after 10,000 conversions. Do this with every test. As long as there are no obvious interaction effects, you can run multiple tests in parallel.
If you wish, you can replace 10,000 with N where N is however many conversions you expect to get in a month. Replace 100 by the square root of N. Real differences below 2/sqrt(N) are basically coin flips. Even after many tests, you are very unlikely to have ever made an error as big as 5/sqrt(N).
If you want to do something more sophisticated than this, you need competent people. But this is worlds better than what most companies do. (And yes, it is better than what Optimizely does for you by default.)
Coming up with ideas to test tends to be fairly easy if you've got a competent product person.
Perhaps I'm biased. I've built multiple A/B testing systems for multiple companies, and know what to do. However it didn't seem hard for me the first time either.
See also: http://www.evanmiller.org/how-not-to-run-an-ab-test.html
See http://www.evanmiller.org/sequential-ab-testing.html for a very similar methodology presented by the very person you are citing. The difference is that I am aiming to always produce an answer from the test, even if it is a statistical coin flip, while he aims to only produce an answer roughly 5% of the time if there is no difference.
If you go with whichever is ahead, you'll reliably produce correct answers for wins at that you can't reliably measure. Yes, there will be mistakes, but the mistakes are guaranteed to be pretty small.
If you insist on a higher standard, you'll learn which choices you were confident of..but now need to figure out what to do when the statistics didn't give you a final answer.
I think that the first choice is more useful. Evan prefers to clearly distinguish your coin flips from solid results. But in the end you have to make a decision, and it isn't material what decision you make if the conversion rates are close.
If you have any sort of decent dashboarding the cost of a wrong decision is really not all that bad compared to the cost of being a purist.
That visible "zone of stupidity" is based on how long it takes to make a conversion which has everything to do with how your product works, and nothing to do with how statistics works. There is absolutely no difference between the graph you expect that leads to an accidental wrong decision and one that detects a correct difference - the patterns that you think you see don't mean what your brain will decide they do.
And more importantly, if you stop in 1/4 of the time that you would have been willing to run the test, the potential loss when you are wrong is twice the worst errors that you could make if you put more effort into it.
Have you ever been in an organization that rolled out an innocuous looking change that killed 15% of the business? I have. Over the 10 months it took to find the offending subject line, there was, shall we say, "significant turnover in the executive team".
Math exists for a reason. Either learn it, or believe people who have learned it. Don't substitute pretty pictures that you don't understand, then call it an explanation.
Let's say you want to figure out the unknown bias on two coins. You flip both continuously and plot the percentage of heads you see. Due to law of large numbers these percentages will eventually converge to true probabilities (which is how I am interpreting the graphs in that blog post).
The bad case is if the two coins are actually "flatlined" in the wrong order so as a pattern matching human you mistakenly believe the rates have converged prematurely. I don't know how to work out the math on this but let's say a "flatline" is visually 100 points or so with no significant slope. Then this should be pretty rare right?
If you want to try to understand what is going on, learn the Central Limit Theorem. That will let you know how fast the convergence is to the laws of large numbers. (There are two, the strong and the weak.)
Usually the decision rule is stated in terms of a statistical confidence test. Usually the confidence test is done poorly enough that it doesn't mean what people think it does.
And the stopping procedure isn't actually very arbitrary. You choose N based on the most data that you're willing to collect for this experiment. And stop when you're confident about which version will be ahead at that point.
So this procedure leaves you confident of having the best answer that you will get from the most extensive test that your organization is willing to commit to running. And the cost to the organization of running it is capped at sqrt(N) lost conversions.
It took a lot of work to get down to the super simple version. :-)
If accounting had been invented recently, it would be hard to find someone who knows what to do and it would be hard to get the rest of the company to operate in a way that allows him to do it.
The more companies do this, the easier it well get. The fewer people do this, the more of a competitive edge it is.
You're mostly right in what types of resources you need, however, I'd add that the "competent web designer" should also be competent in jQuery. For a lot of tests, running jQuery DOM manipulations is required/preferable over the visual designer. Those visual designers tend to fail when dealing with dynamic content and have the potential to halt your entire operation. Of course, you may also have teammates (or you yourself) who can do more than one of these things.
1. Management buy-in to always be testing and always be investing in testing
- It's one thing to have the boss wake up tomorrow and decide "We should do some A/B testing", have the team go off and implement a few tests, then when that's done, move onto the next feature-of-the-month. To really make it pay off, A/B testing needs to be done all the time, to the point where it drives priorities and it makes no sense to commit to the next thing until the tests confirm what that thing should be. Which brings me to:
2. A product design culture that integrates the results of A/B testing full circle back into product decisions
- Why even do these tests if results don't feed back into the product? If testing says header bar A's choices perform/monetize better than header bar B, but the designer sticks with header bar B simply because of his artistic "gut sense", then why are you even investing in testing?
1a. Management buy-in to devote resources to clean up complete/unsuccessful tests.
I witnessed first hand the ball of mud resulting in 10 years of test upon test upon test. It's a huge impediment to quick iteration; simple changes are hard and hard changes are impossible. It's especially bad because test code tends to be poorly designed and documented due to the get-it-out-the-door mentality.
This is fairly well trod ground though and for those that want a cheaper approach, here are some free resources that I like. Some of these are my own blog posts but I'm not selling anything so pardon the self promotion:
1. "The Pitfalls Of Bias In A/B Testing" - http://danbirken.com/ab/testing/2014/04/08/pitfalls-of-bias-...
2. "If You Aren't Doing Basic Conversion Optimization, You Probably Should Be" - http://danbirken.com/startups/2015/05/18/conversion-optimiza...
3. "How not to run an A/B test" - http://www.evanmiller.org/how-not-to-run-an-ab-test.html
4. "Evan Miller's sample size calculator" - http://www.evanmiller.org/ab-testing/sample-size.html
5. "ABBA Open Source A/B test calculator" - https://www.thumbtack.com/labs/abba/
(and a bonus one because the improved sign up flow in the OP still had a password prompt)
6. "Password prompts are annoying" - http://danbirken.com/usability/2014/02/12/password-prompts-a...
Shameless plug: Two open source libraries for A/B testing I created and use regularly.
AlephBet[0] is the JS library to write your test in (with multiple backend alternatives)
Gimel[1] is an AWS Lambda backend you can run for near zero cost. Even at scale.
The Gimel (minimalistic) dashboard is using the algorithms / code from Evan Miller and others to do the bayesian statistical analysis.
1. On the difficulty of A/B testing: It is heavily dependent on a lot of different factors in your organization. Developers' ability to quickly generate variant pages is a huge issue: lots of places simply can't make changes to revenue-generating pages. Lots of places deal with politics that keep them from rationally analyzing A/B tests. And c-level support can ram through a lot, of course.
I know that we're all from planet Vulcan and are very hyper-logical about testing, but most orgs don't work that way. 98% of my work involves doing therapy on people who are used to making design decisions by internal debate. Shifting that culture is tremendously difficult in large organizations where lots of smart people have lots of strong opinions.
Put another way: yes, it's easy to run an A/B test. It is hard to ship the revenue-generating design decisions that come from A/B testing.
2. On where to start: I try to research my test ideas so I know that I'm testing the right things, in the right places. Messaging tends to have much higher impact than specific design elements, so you want to make sure you're communicating to people effectively.
With total blank-slate clients, I get your GA install in order (nobody ever has this), and run heat & scroll maps on all the key pages in your funnel. I also email a survey of your existing customers to see if there are any surprising patterns. And with my bigger clients, I even run usability tests (on usertesting.com) and customer interviews (recruiting through ethn.io) to understand their motivations for purchasing.
None of this has anything to do with actual A/B tests – but it has everything to do with making sure you're not stabbing in the dark when you do test.
Finally, agreed with @gk1 on big, drastic changes. Harder to implement (see point about dev time above) but much more likely to bear fruit – and force the org to really question how they're speaking to customers.
3. It's a new article. The A/B Testing Manual was launched yesterday at abtestingmanual.com. Videos will be ready in a month or two, tops. Launch discount expires this Friday.
Great questions. Keep 'em coming!
When and where do you start? If you start too early, you don't have enough traffic to be statistically significant.
In the article, Patrick said that he started with the trial signup form in his 3rd year.
Most of the advice on A/B testing that I've found is understandably aimed at people with existing business/products that can massively benefit from it. Does anyone have any more material about how to get started from the early stages?
TL;DR - Test big, drastic changes instead of fiddling with button colors and headlines. Examples of drastic changes:
- Entirely different homepage with different messaging.
- Change "Features" page to "Benefits" page and change its content accordingly.
- If you're offering a free download or free product, test asking for an email first.
- If your SaaS has a long signup form, test allowing people to jump right in with just an email address.
- If you're sending a robotic "welcome" email, test sending a very short personalized email instead.
I'd also suggest making sure you approach testing by starting with a hypothesis and then testing that hypothesis. By voicing or writing down your hypothesis, you will be more likely to avoid the "stupid test trap", a phrase I use to describe running tests like changing button colors -- what hypothesis is changing button colors trying to prove? That people like red more, "for reasons"?
As mentioned above, each A/B test you run has a serious opportunity cost. I'm fortunate enough to have thousands of people going through my sales funnel daily but, even then, a test will still take several weeks or months to validate and anything approaching a change of 5% or less will really eat away at your ability to test more meaningful things.
My experience is that most tests will come out under 5% differences. But most organizations will have some test that they could run with a greater than 15% win.
If your business scale does not allow you to run tests and detect the wins you hope for, you're better off following other people's best practices than you are trying to discover best practices from running your own tests.
Also note that most businesses have a "conversion funnel" where there are a number of steps from visitor to getting paid. If your business is big enough, you want to focus on getting paid. If your business is too small to get results that way, you should get started with just the first step in the funnel. That's what Patrick did with the trial signup form.
Seriously contemplated A/B testing them as a joke, but erred on the side of sanity.
Of course, Patrick is all over HN, so parts of the article may seem familiar.