The multi-armed bandit problem (2012)
stevehanov.ca
stevehanov.ca
https://www.aaai.org/ocs/index.php/AIIDE/AIIDE13/paper/view/... https://courses.cs.washington.edu/courses/cse599i/18wi/resou... etc.
For what it's worth the MAB algo in the original post looks like Epsilon Greedy. It's probably better to look into Upper Confidence Bound variants like UCB1 that dynamically adjust how much to explore vs exploit, e.g.:
https://towardsdatascience.com/comparing-multi-armed-bandit-...
The main change in AlphaGo was using a deep learning network to encode a value network for fast rollouts and a policy network for move selection (rather than using the UCB rule). They later removed the value network and rollouts entirely, but even AlphaZero uses MCTS.
Also, I work in this field and I will just say that people _do_ behave differently based on traffic source: i.e. users coming from Facebook behave alike, but different than traffic from Reddit who act similarly to each other. If you were running a self-optimizing thing like this it _would_ make sense to split it up by the different traffic sources and handle them separately.
It isn't impossible - bandits are seeing adoption in medical trials to avoid precisely the problem discussed - but the standard experiment design and analysis techniques you learn in a decent college statistics class or introductory statistics text no longer apply. That's one of the beauties of A/B testing: while it does require substantial thought to do well, the basic statistics of the setup are very well-understood at this point.
I’m looking for variants that win. When I find one that wins I look at it and try to add more of the same flavor to the product.
This process works.
Edit: added in forever. Phone dropped some wording I originally had. I think.
By all means, keep doing it if it is working for you. But don't confuse it as good advice. And stay vigilant.
Again, it may be working in your case. Argument to authority can go a long way. Even ad hom attacks often exist due to a "smell" of the person speaking. It is not, however, logically sound and can easily lead to unsupportable positions.
So, take care. And realize that a lot of the damage of poor practices may be tangential. For example, a belief that the real world can not be described by science.
It's easy to underestimate how complex things are, because we only see some superficial aspects of e.g. a user/software interaction model. This flaw is down to how our brains work -- ref "What you see is all there is".
The reason is that websites are a dynamic environment. All things equal, better controlled experiments are great. However, visitor behavior, especially when from dynamic sources (google serps change weekly), changes all the time.
And that's why I prefer MAB over A/B tests -- A/B tests don't adapt to a dynamic system, so we often wish away the changes in the system to use it. Does anyone go back and re-test their biggest wins?
Yes, absolutely! We do research first, then come up with simple, well-controlled tests. Once we have a winner we can either lock it in, but often we continue to research and experiment on the new knowledge we gained. A hefty minority of the tests I implement build on past wins to further flesh out what works and what doesn't with knowledge and the data to back it up.
MAB is best used where you can generate a bunch of variants cheaply and hope for a 30% gain.
Sometimes, yes, but often it's not for the reasons people think. This seems very much in Danny Kahneman land -- we can't trust what our brain is telling us, that this experience will have similar results in this other context.
I recently worked at a place that ran > 20 websites, huge traffic numbers, with around 100 simultaneous tests. We would poll each employee about their guess for the test winner. The results were about on par with everyone randomly guessing.
The key part is 'not always'.
Using A/B tests just to find the better performing version is a valid use case too. And if you have the traffic, multi-armed bandits are a nice way of automating the whole procedure. Even scientifically speaking there is nothing wrong with them. Their biggest issue is that they require a lot of traffic for significant results.
Out of curiosity, what is your hypothesis for explaining this difference in behavior? Would you say it's primarily due to differing contexts in which a link is posted, or differing populations on each platform? Or maybe these both contribute about equally?
Stated another way: would you expect the same individual to behave differently coming from Facebook versus coming from reddit, if they happen to be a user of both?
These could just be weighting factors so instead of a single % chance per option, every time there's a successful interaction the victory is spread across the factors for that option.
New users could be shown the option where the chance for each option is weighted by the factors they match. Is this a valid approach or would it introduce some kind of selection bias?
Occasionally there is pure scientific interest.
But far more frequently, the purpose of the A/B is to optimize the outcome.
This is why Google Analytics has exclusively chosen multi-armed bandit for its its A/B test framework.
If you want to use the results from one test to inform what to test next, then A/B tests are better optimizing for truth.
Most website changes dont make a significant difference in conversion. If you use MAB, then you dont know if the winner is really better or the result of random variation.
No, you know the outcome of MAB is the best.
If it's clearly the best, it converges quickly. If it's less clearly the best, it converges slowly. Either way though, you haven't lost any more conversions that necessary to find that out, or needed to put in non-mathematically fail safes.
If you're worried about the site changing for people between visits from different traffic sources, you can cookie their MAB/test-assignment on the first visit.
Here is an article from 2013 that describes Googles Multi Armed Bandits algo integration into Adsense:
https://www.cmswire.com/cms/customer-experience/google-integ...
The first time I stumbled across the term "Multi Armed Bandit" was when I read Koza's "On the programming of computers by natural selection" in 1992. When I later got involved in e-commerce projects, it was immediately clear to me that this was the way to tackle the involved optimization tasks.
To toot my own horn a bit, I wrote a blog post about Bayesian bandits: https://eigenfoo.xyz/bayesian-bandits/
It shares the same principle of choosing reasonable options most of the time, but allowing variation to keep from getting stuck in a local optimum.
Other examples: momentum in deep learning, tabu search, random search, etc.
The benefit of doing a pure binary A/B test is that the experiment is so simple that as long as you don't break the cardinal rules (you need a fixed experiment size! Most people don't even do this. Also you need to not bias your control/experiment sets) it is easy to get statistically valid results. When doing multivariate optimization using something that is measured (such as user engagement) rather than evaluated, you need a good amount of data for each configuration to evaluate it, or you run the risk of optimizing for random noise. This is true of even the multi-armed bandit problem: if you vary the exploitation strategies over time, then confounding temporal variables (for example, purchases are higher on weekdays because that's when people make business decisions at work) can invalidate the experiment if not controlled for.
https://www.researchgate.net/publication/301935710_Interface...
https://www.microsoft.com/en-us/research/blog/new-perspectiv...
https://podcasts.apple.com/nl/podcast/microsoft-research-pod...
Sure you have to read a bit more to know why it works, but if you write your code well you could plug this in without any extra trouble. It's not like you need a special optimization solver as a dependency.
I personnaly think that it should replace UCB1 as a baseline when trying bandit algorithms.
[0]: https://homes.di.unimi.it/~cesabian/Pubblicazioni/ml-02.pdf
https://towardsdatascience.com/bandits-for-recommender-syste...
As a cancer researcher I read this as a just criticism of my work.
Just out of curiosity... Have you ever purposefully clicked on an ad on the internet? I honestly dont think I ever have.
ps. I mean an outright overt straight up ad, not, for example, some article linked on HN that is a thinly veiled promo piece for something.
This seems like an obvious extension, and something that someone should have worked on given how long this problem has been around, but I've been unable to find anything on it. Any pointers?
Based on just the article it could be a multi armed chef.