Out of the cesspool and into the sewer: A/B testing trap
blog.asmartbear.com
blog.asmartbear.com
I have a eight A/B tests currently running. Here's how long it took to launch each one. Spot the big bang test: ten minutes, fifteen minutes, fifteen minutes, 30 seconds, 15 seconds, an hour, three hours, three weeks.
Small A/B tests also have small risks associated with them. One of those eight touches a single word on my website. It is almost inconceivable that that one-word change could result in a POed customer sending me email. On the other hand, the big bang test can cause (and has caused) customer support issues for me, despite taking a great deal of time to minimize the impact.
In addition, big bang tests are often less conclusive than you want them to be. My current Big Bang test is significant at 95% right now, against the pivot I want to make. I am of two minds: on the one hand, I want to bow to that inevitability. On the other hand, I'm seriously wondering whether it is the pivot causing the disparity or if it is just implementation details of the pivot. In a standard A/B test, the change is the implementation detail. However, given that I had to adjust something like 30 files, I'm wondering if customers are really rejecting the pivot or whether they just think the graphic I made for it is exceptionally hideous. (I started an A/B test for just the graphic and it is, indeed, getting whupped versus what it replaced.)
Since the big bang A/B test frequently changes mountains and drops you at the unoptimized bottom of the new mountain, you're left either trying to do hillclimbing on two hills in parallel (which is OK, as long as you don't mind your engineering team cutting out your intestines and using them to strangle you) or doing the deeply unsatisfying "Well, I really feel better about B, so we're going to hop over there and then start hillclimbing, then pat ourselves on the back and assure ourselves it was the right decision all along."
If growth/success is not measuring up to one's goals, you have two choices. You can A/B test your way up or you can do something dramatic (hire a salesguy, pivot your product or positioning, etc). I think a lot of startups with big dreams and crappy growth dive in with A/B tests when they should be figuring out how to change how people talk/think about their product(s).
You're right about being expensive. The scarcest fuel for most teams is optimism and confidence. One or two failed pivots can be devastating.
+1 on dramatic!
I manage CPC campaigns for clients, and often times I make A/B changes in the keywords/bids/ads and wait 30-60-90 days before measuring and making changes based on the findings. And even with 90 days and tens of thousands of visits, I find that sometimes my scope is too limited and I get nailed by seasonal/yearly changes that the numbers didn't reveal.
A/B testing in 5K chunks works if everything else in the universe remains constant, but here in the real world that's never the case.
The worst is when you make a change, measure results and implement said change, and then something related to the economy (or even Weather) horks your CAC and suddenly the client isn't so thrilled with your amazing A/B improvements.
By implementing parts of this algorithm you might be able to generate a random walk scenario that is local minima resistant.
A very powerful concept that you don't see done very often.
http://en.wikipedia.org/wiki/Simulated_annealing
Start your optimization process with a high "temperature" (leading to big changes), and let it "cool" over subsequent iterations.
See http://visualwebsiteoptimizer.com/split-testing-blog/using-a... as an example of subtle change.
In my opinion, it is entirely misguided to blame local maxima problem on A/B testing. The methodology doesn't dictate what is that you want to test. It is all about having a hypothesis and then collecting data to prove/disprove it. The post looks like it is blames poor results on instruments while wrong experimental design is the true culprit.
It's like Google testing 40-odd shades of blue to see which performs better, while failing to discover that the best color to use is (say) red.