Edit - and this: http://www.stat.columbia.edu/~gelman/research/unpublished/p_...
Simpson’s paradox is about spurious correlations between variables - conversion analysis is pure Bayesian probability.
It shouldn’t be possible to have a group as a whole increase its probability to convert, while having every subgroup decrease its probability to convert - the aggregate has to be an average of the subgroup changes.
Consider the case where iOS users are more likely to convert than Android users, but you currently have very few iOS users. You then A/B test a new design that imitates iOS, but has awful copy. Both iOS and Android users are less likely to convert, but it attracts more iOS users.
The group as a whole has higher conversion because of the demographic shift, but every subgroup has less.
On that topic – what do you do when you observe that in your test results? What's the right way to interpret the data?
There are two issues at play here -- one is that the sample sizes for the segments may not be high enough, the other is that the more segments you look at , the greater the probability for finding a false positive.