Learning Reinforcement Learning, with Code, Exercises, and Solutions
wildml.com
wildml.com
Regarding the post: This seems like a useful resource. When I read many of these papers, a code supplement makes understanding it so much better. I do a lot of research with RNNs with respect to language modeling and implementing various models when I started researching this field was very useful to get a better understanding. I got great feedback on my implementations where people said it helped them understand.
I think the ML results and demo posts are mostly getting upvotes because the results are fascinating and the posts are often excellently written. Image manipulation in particular lends itself towards posts that are appealing both to people casually skimming articles and to people looking for some technical depth.
However, I'd appreciate more comments of course ;)
I do research in ML so I love seeing these posts, but the fact that words like neural networks has become such a buzzword is somewhat disappointing. I remember all the buzz when Swiftkey released their "neural network" keyboard simply because of the name.
I think it has to do with the fact that HN is a very diverse mix of folks (technical, non-technical, working at startup, working at big co, web devs, mobile devs, infra dev, etc). There isn't a concentration of ML-technical folks like you might find on /r/machinelearning (decent technical ML discussion). Instead you have here a mix of futurology speculation and tutorial discussion - lots of introductory material of highly varying quality (since the people upvoting don't really know if it's good), science fiction passed off as credible opinion, highly technical ML posts that get upvoted a lot but no comments, etc.
Nobody thinks SVM is cool anymore :S
Part of this is due to, I think the immense amount of marketing pushed by Big Corp on their recent AI successes.
Calling something cool/hacking because you don't want to take the time to understand the maths is something Trump would do if he was a programmer
I like to read wildml.com and fastml.com blogs, but I would like to find more simple applications that shows real value without using lots of resources. Perhaps there is a subfield of RL where using some kind of proper human intelligence one can hope to beat those giants provided of unlimited computational and financial resources
The Jupyter notebook is included in the GitHub repo[3], and includes a 'scaled down version' that takes ~5mins to train on a MacBook's CPU. There's also a downloadable 'full scale' model that was trained in ~7hours on a Titan X. It plays the game (on average) better than me...
[1] http://blog.mdda.net/ai/2016/06/23/workshop-at-pycon-sg-2016 (has slides, and YouTube link) [2] http://redcatlabs.com/2016-07-30_FifthElephant-DeepLearning-... [3] https://github.com/mdda/deep-learning-workshop : have a look at notebooks/7-Reinforcement-Learning.ipynb
If you can figure out a way of making RL better than polynomial time there is at least a Turing Award for you.
RL for NLP? I would love to know about counter examples, but I'm not aware of a serious project using RL for NLP, let alone 'widely used.' However I do believe RL makes sense for a number of NLP problems.
Either way, well done. I appreciate a collection of the algos from Sutton's book (great book), and in Python.
You'll find similar applications in state-of-the art models for chatbots for example. Though I agree, "widely used" may be somewhat of an overstatement. But it's becoming more common.
On a side note, I actually think RL makes a lot of sense for many NLP problems and it would be super interesting to build a pure RL approach to language modeling or translation. Nobody has managed to do that quite yet.