Appsumo reveals its A/B testing secret: only 1 out of 8 tests produce results
visualwebsiteoptimizer.com
visualwebsiteoptimizer.com
To add insult to injury, half of the remainder were significant... in the wrong direction.
(Edit to add: null results aren't failures, though. What is the Edison quote: you now know one more thing that didn't work.)
Unfortunately, it's difficult to determine the difference between those two cases. If many of your tests are failing to be significant, it could be that you're simply never investing enough to get the power you need: your users are somewhat inflexible to the design changes you're making.
It could also be that many of your changes simply are pretty insignificant and by the time you wait a full year gathering simply infinite numbers of impressions, you'll find it was just a waste of time.
Again, practice statistics with a great deal of self-awareness. They're only meant to inform.
Before people did A/B test, it was generally assumed on Facebook ads that the title was the second most important factor on CTR so people spent a lot of time tuning title copy.
After lots of people did A/B testing it turns out that title copy has almost no impact, people switched their title to Chinese (for a non-chinese speaking audience) without seeing any change. That "no significant difference" means that thousands of man hours can be spent elsewhere rather than copywriting ad titles.
http://untyped.com/untyping/2011/02/11/stop-ab-testing-and-m...
I am familiar with the maths for A/B testing. The figure of approximately 100'000 hits comes from "Controlled experiments on the web: survey and practical guide". As I stated, this assumes 5% conversion rate and a few other things. Here's the quote:
"If, however, you were only looking for 5% change in conversion rate (not revenue), a lower variability OEC [Overall Evaluation Criteria] based on point 3.b can be used. Purchase, a conversion event, is modeled as a Bernoulli trial with p = 0.05 being the probability of a purchase. The standard deviation of a Bernoulli is √p(1− p) and thus you will need less than 122,000 users to achieve the desired power based on 16 ∗ (0.05 · (1−0.05))/(0.05 · 0.05)^2."
(The actual value is 121'600.)
The formula they use is n = 16 σ^2 / Δ^2. σ is the standard deviation, and Δ is the size of the difference. Thus in this problem, Δ = 0.05 and their formula gives n = 16 * (0.05 * (1-0.05)) / 0.05^2 = 304. This is much more in line with what you get from using a 2 sample proportion test (with H_a: p_1 =/= p_2), ~440 in each group [2].
But maybe I misunderstand their formula.
[1] http://exp-platform.com/hippo_long.aspx
[2] http://statpages.org/proppowr.html
edit: fixed Greek letters and added final comment.
If you go to your reference [2] and enter the numbers 0.05, 80, 0.05, 0.0525, 1.0, you'll see that they come up with a sample size of about 122k in each group (so 244k in both together).
(The figure of 304 or 440 is what you would get if you wanted to detect an absolute change of 5% in the conversion rate: going from 5% to 0% or to 10%.)
Fair enough, it would take a very large sample (122K is close enough) to detect a change from 5% to 5.25%. Being concerned about a change that small seems really silly unless 0.0025 * N visitors * revenue per user is a big enough number to be concerned with. I contend it won't be unless either
(1) N visitors is very large or
(2) revenue per user is very large.
If (1) is true, then testing on 122K users is not a big deal. If (2) is true you probably want to have a much more targeted approach, like someone doing sales.
(Rough numbers: suppose you get 1000 visitors per day and convert at 5%, and suppose each conversion is worth $10 to you. Then you're bringing in about $180k/year from them, and a relative change of 5% in that is about $9k. Seems worth doing a modestly-sized A/B test for, but if it takes 4 months then you might reasonably decide to spend your effort elsewhere. (Or, of course, not: the actual cost of doing the test is rather small. But a lot can change over 4 months.)
Knowing that 2 different approaches are approximately the same is a great help in design... It eliminates that 'Which way is better?' worry and lets you design it the way that looks best, instead.