I have a eight A/B tests currently running. Here's how long it took to launch each one. Spot the big bang test: ten minutes, fifteen minutes, fifteen minutes, 30 seconds, 15 seconds, an hour, three hours, three weeks.
Small A/B tests also have small risks associated with them. One of those eight touches a single word on my website. It is almost inconceivable that that one-word change could result in a POed customer sending me email. On the other hand, the big bang test can cause (and has caused) customer support issues for me, despite taking a great deal of time to minimize the impact.
In addition, big bang tests are often less conclusive than you want them to be. My current Big Bang test is significant at 95% right now, against the pivot I want to make. I am of two minds: on the one hand, I want to bow to that inevitability. On the other hand, I'm seriously wondering whether it is the pivot causing the disparity or if it is just implementation details of the pivot. In a standard A/B test, the change is the implementation detail. However, given that I had to adjust something like 30 files, I'm wondering if customers are really rejecting the pivot or whether they just think the graphic I made for it is exceptionally hideous. (I started an A/B test for just the graphic and it is, indeed, getting whupped versus what it replaced.)
Since the big bang A/B test frequently changes mountains and drops you at the unoptimized bottom of the new mountain, you're left either trying to do hillclimbing on two hills in parallel (which is OK, as long as you don't mind your engineering team cutting out your intestines and using them to strangle you) or doing the deeply unsatisfying "Well, I really feel better about B, so we're going to hop over there and then start hillclimbing, then pat ourselves on the back and assure ourselves it was the right decision all along."