Super cool. How long did this take to make?
I kind of wonder if there is some nice analogy to be made here wrt. Kelly Betting vs Bayesian RL. As in, some version of maximising log reward will have higher median performance than Bayesian RL even though on average Bayesian RL is better. By analogy, the discprepancy should come from Bayesian RL doing vastly better in some unlikely string of world trajectories.