I would love to embrace (or at least try to) such a new approach, but it feels like without a PhD in stats it's hard to get it started.
I would love to embrace (or at least try to) such a new approach, but it feels like without a PhD in stats it's hard to get it started.
For Bayesian ML see https://probml.github.io/pml-book/book1.html and https://www.bishopbook.com/
For Bayesian statistics see http://www.stat.columbia.edu/~gelman/book/
There are some caveats though - and I mention these from the experience of running such solutions on a large scale in production. First, BB-MAB can't adapt to context by design. They only look at click/no-click behavior across the population. So, if your population has two distinct segments - youth and elderly - who behave very differently wrt purchases, the BB-MAB won't pick a different winning advt. per group; its blind to these groups.
The solution is to use something like a contextual MAB - which assimilates user features (or whatever you might throw at it) into the MAB. There are simple ways to adapt simple MABs to the contextual setup [2] (in my experience, these can also be effective) but, of course, the literature in this area is wide and deep.
A second caveat is that if the ratio of the size of the pool of advts. to the number of impressions is high, the BB-MAB won't converge or converge to a good optima; the search space is simply too large relative to the data. In cases like this it becomes important to begin with the right Beta priors, instead of the standard recipe of starting with a Beta that looks like a uniform distribution.
But I wanted to say that when I looked A/B testing few years ago I started off with this book.
https://github.com/CamDavidsonPilon/Probabilistic-Programmin...
And somehere down the line he will introduce the Beta/Binomial method for A/B testing.
In my (again humble) understanding the benefit of doing this in the Bayesian way is both that you can actually get an understandable answer, but also that the answer does not have to be the final result you can continue by adding a loss function for example.
Some very basic Bayesian models can go a long way towards making informed decisions for a/b tests