A/B Testing is Expensive
jamiequint.com
jamiequint.com
When you're first starting out, positioning matters. At http://yesgraph.com we've found copy AB tests to produce incredible lift.
For example, people don't want to "invite" contacts. They do want to "email" contacts though. It's the same flow, but a few words triggered massive lift. The reason it is massive is because it is so unoptimized. So it is specifically at the start where such small tests can matter.
Dan Siroke, CEO of Optimizely always, always stresses in his presentations it is vital to test. The presentations on his SlideShare are awesome: http://www.slideshare.net/dsiroker
However, you only have so many "bullets" to shoot at tests with a small audience so you have to be very picky. Sounds like contact inviting is a key viral feature of YesGraph so it makes a lot of sense to optimize there.
I'm not anti A/B testing at all, but anecdotally a lot of people I talk to about this stuff fire their test bullets on the wrong things and end up with not much to show for it.
A better understanding of psychology goes a long way into intuiting where you may actually be able to see gains, and how to go about achieving them. Sometimes this can be small changes producing large gains, although I have seen large changes produce larger gains more often.
Feedback like this should be taken with a grain of salt, since these people are testers, not necessarily like your users in all respects. But it's still really valuable. I've caught numerous errors that test data would not help me understand easily.
Combine remote usability testing through something like usertesting.com with prototyping, and you've got a really rapid way to get feedback on the cheap, even if you don't have enough site visitors to get statistical significance on a reasonable time frame.
I have a site with a few thousand visitors per day. I had a product I was going to release. I was working up to a big product launch, building up the anticipation over my email list and on the site itself. In that case, I couldn't use natural traffic from the site to test the sales page prior to launch. I had to use cold traffic such as adwords to see which version of the page people responded to. I probably spent about $8k on traffic just crafting and A/B testing the sales page.
But it was worth it. It was like a university education in marketing. Marketing can be so counter-intuitive. So many things I expected to work did not work and vice-versa. But once I had the final tested sales page in place it worked and it worked well. I still get about a a 4% conversion. And most importantly, I knew that when I did finally launch, I had a well-tested and solid page that would convert the large initial influx of customers I got from the building up process before the product launch.
- Variation A and B each receive 20 visits
- Variation A receives 10 clicks while variation B receives 5 clicks
- The confidence interval for Variation A is 90%
(Source: https://mixpanel.com/labs/split-test-calculator)
Also, I wrote an article titled "Creating Successful Product Flows" that is very relevant to this post:
https://medium.com/design-startups/c41ffbce49a1But that was only based on my intuition, not math, and I've never seen anyone give a good discussion of whether "90% confidence" is as definitive as it sounds in the context of a very small sample.
A small sample has less statistical 'power' to identify significant differences where they exist. Put another way, a large sample is more likely to give a true significant result than a small sample.
But, if you do see 10% significance(/90% confidence) in a small sample, this is just as good as 10% significance in a large sample. Although the cutoff point will be more rough in a smaller sample, it's a good standard practice to round conservatively to account for this.
10% is unlikely to be considered a good result for statistics in either case - you can engineer a result by doing 10 tests on nothing and there's a danger you would have unknowingly or unconsciously done this, maybe (for example) by not deciding the sample size in advance. However, there's also presumably strong enough evidence against a harmful difference that you aren't likely to lose anything by following these results.
It can be good idea to do numerous small investigative tests as justification for bigger tests - relying on lots of small tests alone requires consideration for multiple testing (e.g. Bonferroni correction).
The sample has to represent the population, that's fundamental. If the sample is so small that it can't characterise the population distribution, then you have a problem anyway. If you're measuring a events that happen 1% of the time (or 99% of the time), a sample of 100 is not nearly enough.
If you chose an appropriate non-parametric test to cover an unknown distribution with a small sample, it maybe would have zero power (impossible to give a significant result)
The larger the effect size, the smaller your sample size can be before you reach that conclusion.
Most folks don't fix the desired effect size and instead just create a bunch of variants, start the A/B test, wait for the A/B testing framework to shout "statistically significant!", and then declare a winning variant. If the sample size seems "too small" they might not feel comfortable declaring a winner, so they perfunctorily "get a few more samples." Neither of these are rigorous, so it's a bit pointless to debate about which one is "better."
I think the difficulty in reaching 90% confidence is in designing a challenger that is THAT much better the original (i.e. 10 vs 5). Most split tests are shots in the dark. You'll basically need a design or copy that is doing pretty bad and an a challenger that is a lot better (but not obviously good enough that you use it in the first place).
Its fair to A/B test things you expect to produce high leverage changes. That was actually part of the point of the article, no small tests. Focus here first, consumer psych helps you figure out where these opportunities are.
Once you get through these big opportunities though even respectable gains (e.g. 10%) take a lot of traffic to measure. For example, seeing a 10% gain in a 50% conversion rate takes around 2500-3000 visits to A/B test at 99% confidence. Seeing a 10% gain in a 10% conversion rate at 99% confidence takes 10 times more traffic than that.
Why? Why are you so worried about controlling false positives that you're willing to eat a whole bunch of false negatives?*
You're not administering expensive drugs to cancer patients, you're designing a website! If you mistakenly think that green buttons perform better than blue buttons when the actual truth is the null hypothesis that they perform the same, that's not the end of the world.
* and I do mean a whole bunch; in that scenario, moving from alpha=10% to alpha=1% means you increase your false negatives by something like 3x. The power calculations:
R> power.prop.test(n=20, p1=0.5, p2=0.25, sig.level=0.10)
...
power = 0.4951
...
R>
R> power.prop.test(n=20, p1=0.5, p2=0.25, sig.level=0.01)
...
power = 0.1646
...
R>
R> 0.4951/0.1646
[1] 3.008And even if you do get lucky and get a test like the one you described above, chances are, you want to continue to revise the page and make more subtle changes which will mean you need a much larger sample size even to reach the low bar of 90% confidence.
Yes, you'll need moderate levels of traffic for split tests to be effective, so if you don't have the traffic or time to wait around, you should be talking to your users.
You just plug in
1. the number of pageviews your page got in the last month
2. the number of conversions that resulted from those pageviews