1. Most multi-armed bandit algorithms assume that the potential reward for each lever is the same each time you pull it. Unfortunately web traffic does not look like this - there are daily, weekly, and monthly cycles in conversion characteristics, with large random fluctuations on top. An A/B test can ignore this - who got the better traffic is just another random factor that comes out in the statistics. How much would this impact a multi-armed bandit approach?
2. My understanding is that multi-armed bandit algorithms assume that feedback is instantaneous - pull the lever and get the answer. But this is often not true. You send a whole batch of emails before getting feedback on the first. Depending on the business, incoming users can take time to convert to paying customer. I've seen places where the average time to do so was weeks. What decision should be made in that time period of uncertainty? Worse yet, what if one version speeds up conversions relative to the other? I know how to tweak an A/B test to handle this issue (just wait then look at cohorts that should have converted under either), I don't know how to modify a multi-armed bandit algorithm to do so.
3. As a practical matter, companies don't want to keep tests going indefinitely. There is a real technical cost to maintaining a complicated mix of possible pages that can shown. You want losers to be removed from your code base. A/B testing is well-suited to doing that. A multi-armed bandit approach does not.
4. Most companies don't even do A/B testing correctly. I fear that pushing a more complex scheme makes them less likely to try it, and increases the odds of mistakes.