I had the basic pre-college understanding of probability and statistics when I started as the senior architect at a company producing multi-variate testing software.[1] When you dip your toe into that pool, Taguchi is the first thing one reads about, so the team implemented that.
It became apparent that Taguchi wasn't really appropriate or sufficient for web-based testing, so the team learned, devised, and implemented more appropriate MVT models.[2]
One notable bug that we discovered involved a self-optimizing test. The idea was that, once we reached a certain confidence level, we would slowly grow the number of targets that were fed the most successful variant.
We had a minor (on the order off off-by-one or switching a < and a <=) code error that grew the successful variant too quickly, at a point where the confidence level was effectively non-actionable.
As I recall, it took us about six months to notice, and none of our clients noticed.
MVT, and especially our implementation, is obviously much more complicated than straightforward A|B testing. Given the fact that no one was able to sniff out such an obvious error when their tests didn't improve conversion as much as expected has left me with the idea that, while testing is not snake oil, I have 99% confidence the population involved in split testing has only a superficial idea of what they're doing.
[1] I had previously implemented a very simple Apache plugin, mod_gating, that I should clean up and throw on github. Most of the work was in the lexer for the configuration file. :-)
[2] Much of the design of appropriate statistical models was done through consulting with statistics departments at a couple local top-ten universities. We figured advanced stats is like cryptography, if you're not an expert in the general field and you come up with a "proprietary" solution, you're probably screwing something up.