Confidence.js – make sense of your A/B test results
github.com
github.com
Every startup constantly gets told to A/B test everything, but nobody ever tells them that startups rarely have enough initial traffic to do any meaningful testing. As a result you end up spending a lot of work on what is effectively a random landing page or email template that might actually be worse than what you started with.
The other "gotcha" is optimizing for a short term metric (converting visitor to trial) but failing to track whether that cohort actually converts as well to long term paying customers.
Evan Miller has a great sample size calculator to illustrate this: http://www.evanmiller.org/ab-testing/sample-size.html
In the example above, Google's MDE is 0.005% and the small startup's MDE is 60%. If the small startup's baseline conversion rate is 5%, then with 95% confidence and 80% statistical power, they can reliably detect a 60% effect (positive or negative) with 1,968 visits.
It's not just the 60% increase that you should be mindful of. I bet you all small startups want to know if the change they made to their homepage decreases the conversion rate by 60% or more. With an a/b test, they can do that.
Another point to keep in mind is that you can learn a lot from a statistical tie (not enough data to conclude there is a difference). In fact, no matter what your traffic levels are, most of your experiments will be a statistical tie. It's important to learn what doesn't work just as it's important to know what does work. This can really help you with prioritization. In the example above, the small startup can use the statistical tie to draw the conclusion that some tasks in their product roadmap will result in less than a 60% effect in the conversion rate and can be prioritized as such.
But ultimately it will be quite difficult for the small startup to produce variations that have at least a 60% effect in the conversion rate, but it's not unheard of and again, you can learn from a statistical tie.
Just because you have low levels of traffic doesn't mean you can't learn anything from a/b testing.
Some people reply to this by saying "Well then we'll just ship it and measure results without testing." That is a test... but it's poorly designed. If you can't see the improvement in a controlled experiment, it's rare you'd be able to see noticeable improvement in before/after.
Confidence.js is based on the A/B testing code that we use at sendwithus.
A/B testing math is hard and we've worked really hard on making this great. Let us know what you think!
Thanks for bringing that to my attention - I agree that our explanation is a little more convoluted than it needs to be. I'll update that later today to be more "readable" ;)
This is where I personally think javascript programming is going wrong at the moment, so much boiler plate.
You're also programming a bit too defensively. Ask yourself, do you honestly need all those checks? What would actually happen if you got rid of the checks? I think your code would be clearer without them and the programmer would still realize he'd made a mistake and passed in nulls.
Confidence.js comes from how we operate on A/B tests at sendwithus - we think of them as variants associated with results. While the code could definitely boil down to a few simple math functions, we find objects to be a good way to organize our data models.
Your project is great, I'd totally use at work. Also, it was to learn the math behind that. Thanks!
Can you recommend a resource that explains the Math for people who slept through statistics class?
I actually need to brush up on this topic for a new project so anything you can recommend would be greatly appreciated.
Thanks!
We did a lot of digging and there aren't a lot of resources that are easy to digest - but this one really stood out to us:
http://visualwebsiteoptimizer.com/split-testing-blog/what-yo...
I'd love to write a summary of everything I've learned - maybe a blog post in the near future!