A/B testing in Mailchimp: 7 years of successful experiments
blog.mailchimp.com
blog.mailchimp.com
At the time at least, they had no proper way to segment cohorts, meaning that in order to control for things like length or circumstances of subscription, we had to do multiple smaller sends to lists that were subsets of our subscribers and aggregate the testing data afterwards, rather than being able to do just one send. On top of this creating a lot of extra work setting up the sends, it skews the statistics on things like SPAM complaints.
What caused the account to be banned was that some of the smaller lists would occasionally see spikes in SPAM complaints. These spikes were well within statistical norms. However, because MailChimp will ban an account that breaks a threshold of complaints for a single list in a single sending, we triggered the ban, despite the list and send as a whole, being well within their limits. Had the AB testing framework been built properly, to account for cohorts, we would have had no issues.
I tried to speak with their CS to explain the issue, but it seems either they could not understand or weren't interested. All of this, after multiple communications with their tech team about the limitations of their frameworks and how they could improve their AB testing.
If you are looking to do anything beyond incredibly basic AB Testing, I would look elsewhere. We ended up at SendGrid, though my recent issues with their billing prevent me from recommending them either.
Would say more but I don't have time.
This definitely effects the power of tests. However there is an early stopping aspect in the decision to wait longer to collect more information, or to stop the test. This is analogous to deciding to collect more samples or stop a test, which is classic early stopping. In both cases you're making a decision to stop/continue without considering its impact on your error rate.
On closer reading I note that MailChimp by default avoids this issue, by making a decision after a fixed period of time. However they aren't controlling for loss of power and so forth.
-----
Other points:
- With a finite population, such as a mailing list, minimising regret is a better objective than statistical significance. (Want to know more? Sign up at bandits.mynaweb.com and it will be covered in due course.)
- There is a multiple testing aspect to what MailChimp does, and I bet they don't control for false discovery rate.
- Also some post-hoc reasoning.
- Statistics is hard, and it's easy to criticise others and hard to do well oneself. I hope this post will be taken as constructive criticism.
A/B testing may be better to figure out one specific aspect (time to send, sender name) but wouldn't multivariate testing make sense if you were looking for the best combination of those variables? i.e. if I wanted to find out that sending an email at 9pm, from a corporate account, with larger images is the best combination?
I'm interested to hear others' opinions on the pros/cons.
Testing one difference guarantees that one side will beat the other (whether or not there is a true difference).
Multiple testing has a danger of the results appearing inconsistent/random (again, whether or not there is a true difference).
Encouraging people to test one difference benefits MailChimp, for exactly the reason that they say - the results will always appear to be relevant.
One easy experiment is to do the null A/B test. Send two sets of identical mails and analyze the difference of the result of the two sets. Usually you will get more variation that most people naively expect.
Today, what are my options if I just want to send HTML emails through outlook to my 3500 subscribers, without being flagged as SPAM ?
Mailchimp in particular is pretty aggressive about terminating your account if too many people mark you as spam.
http://help.mandrill.com/entries/22242948-What-is-Mandrill-
> Technically, though, you can send any legal, non-spam emails through Mandrill.