OP here. Thanks for making this connection! I hadn't heard of OFU before, but it sounds interesting. I'm actually working on another post about "optimizing around the optimal" and use multi-armed bandit as an example =). As Jeff Atwood mentions in his post on SOWH, writing about this has been a great way for me to think about the topic. If it turns out that smarter/more experienced people have already thought about/written about this subject at length, I'm not at all disappointed that I'm retracing their steps. In fact, it means that I independently found the right path!
I see this Coursera course on OFU: https://www.coursera.org/lecture/practical-rl/optimism-in-fa... and what looks like Auer's original UCRL paper http://papers.nips.cc/paper/3052-logarithmic-online-regret-b... Do you have a preferred source for learning about this? Thanks and Merry Christmas!