67 karma · joined February 17, 2017
This is seriously cool BTW.
I kind of wonder if there is some nice analogy to be made here wrt. Kelly Betting vs Bayesian RL. As in, some version of maximising log reward will have higher median performance than Bayesian RL even though on average Bayesian RL is better. By analogy, the discprepancy should come from Bayesian RL doing vastly better in some unlikely string of world trajectories.
Also, thanks for the positive re-inforcement idea.
Its quite satisfying you get the property from the rotational aspect of the QM.
Then again, AMD's making pretty bold claims about their pricing for next year, so you could hang on.
Do you think that chaotic systems would be better analysed by quantum computing over classical computing? More generally, is BQP powerful enough to deal with chaotic systems in the same way P is for linear systems?
Oh, and I think your review of "Enlightenment Now" was a bit too rosy. When he analysed the data he's superb, but he seems to lambast people he heavily disagrees with. Its a tad disheartening.
The AI is rewarded if at each checkpoint the state vector its produced is sufficiently aligned with the videos.
I guess that's the initial training to deal with sparse rewards.