1. We assume that the distribution of rewards for each lever is fixed (over the short term). This allows the reward to vary randomly so long as the average reward (over the short term -- days to weeks) is constant. There are more complex schemes, which allow for greater variation in reward, but the initial Myna offering is intended to be directly comparable to A/B testing.
2. It's not necessary to assume that feedback is instantaneous. Basically you can continue to make suggestions (pull levers) in proportion to your best estimate of their expected return and the maths holds. Very long conversion cycles will cause problems for any system, I think, as you'll spend a long time in a random walk. In these cases we recommend using a proxy measure which is correlated with conversion, if one is available. As for one option speeding up conversion, I don't think that will matter as you'll simply refine your estimate of one lever faster. I haven't thought too much about this particular issue; it would be worth doing some simulations to see.
3. Just turn off the bandit when you've satisfied with the results. That is, you can use Myna like A/B testing, and the Myna UI displays confidence bounds for this very reason, while still getting the benefits of optimising as data arrives.
4. I actually like bandit algs as there is less for the user to mess up. You don't have to worry about how much data to collect, what p value to use, and so on. Just set it running and it optimises automatically.