DarkForest: Deep learning engine for playing Go from Facebook research
github.com
github.com
This is perhaps the best of a hundred scifi novels I've read in a few years. These novels will make you question the wisdom of sending traceable radio signals into space. Stay quiet, it's a dangerous Dark Forest out there.
Not to mention that with the paperback I own something, while with the kindle version I have a revokable license to download a DRM covered file...
My problem is how many Kindle books I see are $9 in the US store and then $30 in the Canadian store. For the same digital content :/
Unfortunately just having fascinating premise and good story is not good enough for me. I'd almost rather it be an essay, a thought experiment, than a book.
Perhaps Facebook will start to catch up, or Microsoft will suddenly appear with something startling, but for the moment the only real competition for AlphaGo comes from top human players. (Perhaps not even them? I don't think anyone really knows.)
Sounds like EXACTLY the motivation a contest would help to foster.
And because we're doing silly memes, to accept the challenge, Microsoft and Facebook would probably have to go into the competition with this kind of attitude: https://www.youtube.com/watch?v=9ZYg4ZbcOPQ
Maybe someone can clarify but it doesn't seem like hoarding this model (specific to playing Go) would help Google have an advantage in their search algorithms over Facebook, for example.
Both do the cool thing of combining Monte Carlo Tree Search (which used to be state of the art, and is like really smart brute force tree search) with Deep Learning, and there is not much more to DarkForest than that so we can describe that first. The Deep Learning portions involves training a Deep policy Convolutional Neural Net, meaning a neural net that takes in a Go position (actually both take in Go-specific features about the position) and outputs the best moves. Datasets from human moves are used to get this, so its pretty easy. This policy neural net is used to guide the Monte Carlo Tree Search, which just means using the net's move predictions to play out a bunch of games into the future to evaluate the best move. In classic MCTS you play out to the end of the game and estimate the 'value' of a move by just ratio of wins it gets at the bottom of the tree, but the estimation of value is a more complicated for both systems. In DarkForest they 'Use PUCT and virtual loss. Remove win rate noise'. The move with the highest estimated value based on the tree search (guided by the neural net) is chosen as the next move.
AlphaGo does a bunch in addition to this which is hard to sum up, but basically: 1. They train a fast and a slow (but better) policy net, using the same dataset. Both of these are used to guide the tree search. The slow but better is used less often than the fast net, so a lot of games can be played out in the tree search. 2. The slow but better policy net is improved through reinforcement learning - making the system play itself and learn from that rather than just the human moves dataset. 3. Since there is a policy network that can choose moves, one can reasonably assume we can also get a 'value' network to evaluate how good a position is. This is also done, and the value network is used as part of an equation to compute the value of positions in the tree search (there is also a much simpler hard coded value equation, and the final value used as part of tree search is a weighted combination of these two). 4. All this is run in an absurdly, radically distributed manner with an insane amount of compute power - on the 'non distributed' version AlphaGo uses 40 search threads running on 48 CPUs, with 8 GPUs for neural net computations being done in parallel, and in the 'distributed version' it uses more than a thousand CPUs and close to 200 GPUs. I don't even know if that's up to date, those number are at least a few months old.
To sum up: DarkForest combines a policy network trained by supervised learning to guide Monte Carlo Tree Search for move selection. AlphaGo uses a small/fast policy network trained by supervised learning AND a slow but better policy network trained by both supervised and reinforcement learning to guide monte carlo tree search, and computes the value of final position with a combination of a hardcoded value equation and a 'value' network derived from the slow policy network. AlphaGo is absurdly, radically distribued (which helps a lot). Oh yeah, it should noted the amount of engineering required for AlphaGo involved dozens of people (two dozen are listed on the paper), whereas Darkforest seems to have been mainly developed by two people.
Yeah, probably not given this recent announcement:
https://cloudplatform.googleblog.com/2016/05/Google-supercha...
I suppose the question one could then ask is "will AlphaGo's approach wind-up being emulated over time or is it going to be something like a cul-de-sac?
How many single algorithmic challenges are worth expending this much effort on? Could AlphaGo's approach be applied to other such problems? Will increasing processor speed just make all this effort moot? Is AlphaGo something like Deep Blue (the custom computer that beat Kasparov and then was dismantled rather than being developed further)?
My take is that the approaches of AlphaGo are more applicable to other problems than DeepBlue, but not by much. Rigid rules make tree search and reinforcement learning easily applicable to Go, but not so much for many real life problems. I made a small diagram to illustrate this point (http://www.andreykurenkov.com/writing/images/2016-4-15-a-bri...) as part of a series of posts about Game AI (http://www.andreykurenkov.com/writing/a-brief-history-of-gam...).
Still, the general ideas of supervised learning followed by reinforcement learning, training multiple models of varying complexities from the same dataset, and combining tree search with learned models as they did are useful general ideas. Hybrid methods as a whole will become increasingly common, I think (no doubt self driving cars already are very complicated hybrid models).
I would outline the theory here, but it's kind of a book spoiler. So be wary if you have interest in reading it eventually and want to know what I'm referring to now.
Since Go is zero-sum the only reasonable strategy is to be maximally malicious and assume the other guy is too.
I'm curious if it's for Torch, or perhaps there's something more fundamental about why it's a good language for AI that I don't know yet.
Still, I think ultimately, LUA is a liability for Torch, and it's not helped by the fact that not only is TensorFlow a superior competitor (in my opinion) but also because TensorFlow is based on Python (which is vastly more popular than LUA).
Despite all their investment in Torch, it wouldn't surprise me if eventually, Facebook transitions to TensorFlow because it's probably what it will take for them to effectively compete against Google on the machine learning front.
https://m.facebook.com/story.php?story_fbid=1015344288418214...
[0] https://cloudplatform.googleblog.com/2016/05/Google-supercha...
It's not even the strongest bot on KGS.
It seems like an indication that either they haven't incorporated the advances of AlphaGo or that Alpha succeeded through the investment of a huge amount of tuning time and processing power rather than through specific advances.
The version of AlphaGo that beat Fan Hui won 77% of games against the strong engine Crazy Stone giving a 4-stone handicap, even when running on a single machine. Crazy Stone has improved since then, though.
Translation: Google beat us and there's no point keeping this private anymore. ;)