Announcing Evan's Awesome A/B Tools
evanmiller.org
evanmiller.org
As an alternative to the Chi-squared calculator, people might want to check out ABBA, a tool I wrote here at Thumbtack:
http://www.thumbtack.com/labs/abba/
It shares the visual component and the linkability, two great features you've nailed. It lacks the live updating and the slider, which is really cool and something I've wanted to add to ABBA for a long time. On the other hand, it supports multiple groups compared against the baseline simultaneously and incorporates a correction for multiple testing into its p-values and confidence intervals, which can be handy. It also uses different mathematics under the hood, but that's not going to be a concern for most users.
Glad to see another step towards a more statistically-aware world!
(For example, with an 8% baseline conversion rate, a 1% absolute detectable different, 85% power and 10% significance level, your tool says 10,583 per branch. R's `power.prop.test` gives sample sizes of 11,182 for a positive change (8% vs 9%) and 9,974 for a negative change (8% vs 7%). The exact midpoint is 10,578.)
I even built a bookmarklet for quickly grabbing numbers off off of a page (usually for Google Analytics data) and passing them to your calculator. https://github.com/yahelc/ABBA-bookmarklet
http://www.experimentcalculator.com/
*edit: yours is awesome. Nice work.
My goto reference is still the wonderful btilly presentation about how to A/B test properly, with nice examples and code snippets: http://elem.com/~btilly/effective-ab-testing/
He provided a full on javascript tool that isn't as polished but works great: http://elem.com/~btilly/effective-ab-testing/g-test-calculat...
Not to belittle this nice package -- it looks like a basic stats calculator for calculating sample size confidence levels with friendly visualization, and I'm just trying to understand what's being valued on the market/industry right now. Is it because current A/B testing software doesn't provide these basic calculations? Or is it that it's well presented and visualized to a lay crowd?
The statistics being used in the A/B testing world is stuff you would've learned in your very first statistics class, it seems. From the success of Optimizely and VWO that the focus is definitely more on the viz and presentation than it is on using any cutting-edge techniques.
Good to know there's plenty of opportunity to bring better stats to high tech. Of course, I understand a lot of the value comes from making those things applicable and meaningful to the users...
As for bringing better stats to high tech, I've thought of this as a wonderful challenge. I'd especially like to see more focus on not violating modeling assumptions (more non and semi-parametrics), and using some more modern techniques from the ML and Bayes literature.
Hypothesis testing is so last century :). Would love to discuss it further with some similarly-inclined HN folks.
Sorry for all the parentheticals. You'd think I was a lisp programmer with the amount of parenthesis I used.
Is there a way to view the source code or formulas you use on your pages? There's been a strong push in the academic statistics world for reproducible research, which means public data, open source statistical code.
I ask because I'm curious about your two-sample t-test. Does it pool the variances for all values of the two standard deviations? Pooling doesn't make sense when one sd is 50 and the other is 2...
I assume that most people have a conversion rate of X, say 30%, and want to increase this by Y, say 20% (30% to 36%). If I consider the type of headline many blog posts and reports have, they are like "How we increased sales, trials, whatever by 50%". That's how they think.
And well, I usually aim for 95% significance.
We met a few times during last years Gig Tank (I am one of the cofounders of http://banyan.co). Awesome to see you killing it. Are you planning on coming back to Chattanooga anytime soon? Would love to grab beers. My email is in my profile, and I would love to reconnect.
https://mixpanel.com/labs/split-test-calculator
They use two-proportion z-test.
I would love to see Optimizely and VWO embrace similarly non-ambiguous and functional reporting as a default.
EG - just introducing Chi-Squared testing into a discussion with clients or teams that think that they're A/B testing properly by following Optimizely's graphs usually turns the discussion on it's head - "you mean there's a RANGE? well how can we be certain?" etc.
Great work, thank you!
I guess the question I'm boiling this down to is... Why are graphs and comparisons of results that Optimizely or VWO produces not good enough?
I've seen Optimizely call something a "winner" with 95% confidence after 48hrs.
The triangular method we use w/the off-the-shelfers is something like:
a) Optimizely base stats
b) Convergence point analysis (useful to correct for day-of-week / unique traffic swings)
c) Chi-Squared Testing which provides a range so that you can actually assess the risk of a high-confidence test. eg look at the example in evan's tool which shows 8.5%-22% and 13.%-28.9%. This means that Sample 1 could be as HIGH as 22.1% conversion and Sample 2 could be as LOW as 13.3%. If this was rated as a high-confidence test that Sample 2 was rated higher than Sample 1 you could be potentially risking a signifigant conversion decrease if you went with Sample 2. EG needs more data and don't just buy into the "this one is better"
I suspect they removed them for the same reason I tend to avoid the subject when discussing testing with non-technical people: nearly everyone is numerically illiterate, and looking for the "easy" answer. They want a used-car salesman, not a mathematics professor (i.e. "no more of this 'confidence' gobbeldygook -- give me the bottom line").
Sad, but in my experience, nearly universally true.