Reinforcement Learning: From Zero to State of the Art with Pytorch 4
github.com
github.com
Once you've verified that the implementations are correct, it is easier to start the journey to reproduce SOTA on harder problems by playing around with the side-tricks that are often employed.
Pytorch 1 is not available yet.
I suppose it means Pytorch 0.4. but it doesn't sounds so good. We should leave marketing outside technical explanations.
What will happen in the future when there is a Pytorch 4 and somebody find this repo?
Well, if it was changed because of my comment, I'm happy to be useful. If it was not changed and I dreamed it, my excuses.
This new one seems to not mention Q-learning, is that because all these examples are based implicitly on Q-learning, or are these totally new alternatives?
You might think you can fix this by making the action an input to the Q network and keeping only one output; then you could find the action with the highest output. But due to the nonlinearity in the neural network, this is an intractable nonconvex optimization problem.
So instead, you train a neural network to output the action given the state. The algorithms are harder to understand, because Q learning is kind of like supervised learning but policy gradients really aren't. A lot of algorithms (A2C, DDPG, TRPO, etc.) still use one-output Q network (as described in the previous paragraph), but this is just a part of the learning algorithm * . Once training is done, you throw away this Q network. The learned behavior is entirely contained in the policy network. These methods are usually called policy gradient methods.
This article covers policy gradient methods only.
* it's possible to do "pure" policy gradients using only the empirical return, but the Q network helps reduce the variance of the gradient estimate and stabilize the learning.
If you can go from zero to SOTA from a web tutorial in anything less than, I don't know, a year, then the field is severely underdeveloped.
But let that be an opportunity: what it really means is that the field is quickly growing and there aren't well-established experts or leaders.
I think someone with a general graduate-level background in most fields of ML or computer science could acquire a working knowledge of the SOTA in a particular associated sub-topic after reading through a half dozen or so research papers and code. That doesn't seem unreasonable or surprising.
This is not the spirit that some of us want to see in HN. Efforts are always appreciated. The field is improving incrementally, in a very fast way.