Evolution Strategies as a Scalable Alternative to Reinforcement Learning
blog.openai.com
blog.openai.com
Deep Forest: https://arxiv.org/abs/1702.08835
http://www.cv-foundation.org/openaccess/content_iccv_2015/pa...
That's Deep Neural Decision Forests; they benchmark against MNIST and obtain state of the art results.
The other paper is about improving decision forests, although it also uses MNIST as a benchmark.
Apologies if I'm wrong and you did mean the Zhou & Feng paper.
I also agree that vanilla MNIST is pretty useless for Computer Vision researches and was trying (awkwardly) to support that idea by showing how these other non Deep Learning techniques performed equally well.
Variants of ES have been used for years (http://dl.acm.org/citation.cfm?id=1645634). The article seems to ignore almost all the work in robotics, e.g. from Jan Peters' research group (http://www.jan-peters.net/ -> publications).
The good thing is that we have one more paper that justifies this research direction and a little bit more public attention.
I had a similar example in my 20 year old book 'C++ Power Paradigms' in which I used a genetic algorithm to train the weights in a recurrent network. As a performance hack, weights were initially represented by just a few bits, and the bit length would gradually be increased, which greatly increased the search space. I never got this to scale past small networks, but I have thought of revisiting my old code since I have a lot more computing power available now, compared to 20 years ago.
Temporal credit assignment is another problem: the policy is updated after a full episode and there is no way to use the information which step was responsible for which reward. Policy search usually works well if the value function is very complex and the optimal policy is simple.
So using 80 times more machines makes you 60 times faster (assuming those are the same machines) "while performing better on 23 games tested, and worse on 28"[0]?
[0] The paper for this blog post https://arxiv.org/pdf/1703.03864.pdf
Asynchronous advantage actor critic (A3C) https://arxiv.org/pdf/1602.01783.pdf
In the end, it doesn't matter that much which approach is taken because it's all classification problems. We just need the solution matrix, and ideally what computation went into solving it. I feel that this simple fact is lost amidst the complexity of how ML is taught today.
ML isn’t accelerating because of better code or research breakthroughs either. It’s happening because the big CPU manufacturers didn’t do anything for 20 years and GPU manufactures had their lunch. ML is straightforward, even trivial in some cases with effectively unlimited cores and bandwidth. We’re just rediscovering parallelization algorithms that were well known in functional programming generations ago. These discoveries are inevitable in a suitable playground.
I used to have this fantasy that I would get ahead of the curve enough to be able to dabble in the last human endeavor but I'm beginning to realize that that's probably never going to happen. Machines will soon beat humans in pretty much every category, and not because someone figures out how to make it all work, but because there simply isn't enough time to stop it now. There are a dozen teams around the world racing to solve any problem and anyone’s odds of being first are perhaps 10% at best. Compounded with darwinian capitalism, the risk/reward equation is headed towards infinity so fast that it’s looking like the smartest move is not to play.
Barring a dystopian future or cataclysm, I give us 10 years, certainly no more than 20, before computers can do anything people can do, at least economically. And the really eerie thing is that that won’t be the most impressive thing happening, because kids will know it’s all just hill climbing and throwing hardware at problems. It will be all the other associated technologies that come about as people abandon the old hard ways of doing things.
2. Tie future awards to action at this timestep...
Can anyone help with better nitty-gritty explanation ?
The behavior of agents is determined by a "policy function". This function takes in inputs (e.g. what the agent sees) and outputs actions (e.g. what the agent does). The policy function has a set of internal parameters that determines the precise mapping from inputs to outputs.
In their work, they used a neural network as the policy function. The parameters are just all the weights of the network.
In a simple version, you start with some random weights for the NN. Then you make many copies of the network, each with a slight random variation made to the weights. For each of these altered networks, you use them to control an agent for a while, and see how well the agent performs during that trial period. Based on how well the different variations do during their trial runs, you adjust the weights of the network a small amount. You adjust the weights to be more similar to the variations that did well. Then you repeat the process indefinitely (generate new variations, test them, etc.).
[0] http://stackoverflow.com/questions/7787232/difference-betwee...