A Bayesian Perspective on Q-Learning
brandinho.github.io
brandinho.github.io
For instance, you couldn't drop the agent into a different environment with the same rules (avoid red squares, find the green one) and expect it to do anything.
Obviously you can change this with a different representation of your state space, but that's a completely different problem, and much harder.
Is there any interesting work you could point to in that space? My expertise is in statistics in general, not really Q-learning or reinforcement learning.
I kind of wonder if there is some nice analogy to be made here wrt. Kelly Betting vs Bayesian RL. As in, some version of maximising log reward will have higher median performance than Bayesian RL even though on average Bayesian RL is better. By analogy, the discprepancy should come from Bayesian RL doing vastly better in some unlikely string of world trajectories.