MuZero: Mastering Go, chess, shogi and Atari without rules
deepmind.com
deepmind.com
> "For the first time, we actually have a system which is able to build its own understanding of how the world works, and use that understanding to do this kind of sophisticated look-ahead planning that you've previously seen for games like chess.
> "[It] can start from nothing, and just through trial and error both discover the rules of the world and use those rules to achieve kind of superhuman performance."
> [...] MuZero is effectively able to squeeze out more insight from less data than had been possible before, explained Dr Silver.
https://www.bbc.com/news/technology-55403473
It seems like we're getting much closer to artificial general intelligence from two directions: reinforcement learning (such as MuZero), and sequence prediction (such as GPT-3 and iGPT). Very interesting times to be in the AI field.
The one exception is self driving. I listened to the Lex Fridman interview with ceo of waymo recently and he made a case for the controlled environment (e.g. separate detection from decision making and planning) and pushed back against the end to end approach that doesn't make any preconceived assumptions about the environment. As an example he takes red lights. They're clearly human engineered signals, so it makes sense to have a module that can explicitly determine the signal as opposed to learning the behavior
But that's true about other games as well and end to end methods still outperform. Which makes me ask, is end to end learning an inevitability for self driving as well or is this the one domain special due to complexity or other aspects?
In order for a machine to learn driving the same way it learns Atari games, it seems that it would need an extremely high fidelity virtual environment to learn in. The high fidelity requirement would necessitate a lot of up-front investment in trying to get the simulations right. You might spend a whole career just trying to build a drivable Virtual Philadelphia as challenging as the real thing. The details would also make it much more expensive to run training sessions at high multiples of real time.
Given those factors, I'm not surprised that self-driving vehicle experiments just use real environments and don't try to learn the fundamental rules from scratch. But it's an interesting point that these choices may make it harder for agents to keep improving.
I wonder if Google would be willing to pay people to add cameras to their cars to collect real world data in far larger scale.
Whether this will work remains very much an open question.
And the coolest thing is that this could also be used as a basis for a AAA video game, and these tend to make billions as well these days, so it's a win win for everyone. AI companies with funding should invest heavily into virtual reality and gaming, because they will need to perfect this to train their models.
It wouldn't translate to the real world; unless maybe you add enough noise to the simulation to prevent the AI from using too simple cues, but at that point, it's questionable how much insight the AI could still distill from the simulation.
Without any concept of the rules you have no way of even knowing that you’ve set up the pieces for a legal starting position, never mind executing a legal move to open the game.
This is really bizarre.
Basically, AlphaZero was provided with a simulator that was able to distinguish legal and illegal moves and determine which future game states would be wins or losses. This was used to generate the search tree of possible states and actions.
MuZero doesn't have access to a simulator, it only has access to its direct environment. MuZero excludes actions that are immediately illegal, which solves the problem you mention in your penultimate paragraph, but it needs to learn the game's dynamics in order to determine which future moves and states are possible.
> AlphaZero used the set of legal actions obtained from the simulator to mask the policy network at interior nodes. MuZero does not perform any masking within the search tree, but only masks legal actions at the root of the search tree where the set of available actions is directly observed. The policy network rapidly learns to exclude actions that are unavailable, simply because they are never selected.
MuZero still masks legal moves, but only at the root. All its parts are eventually trained on the output of its root, and so learn the legal moves.
The justify this root level masking by how the Atari will only allow you to perform legal moves, while a weak enough player may consider illegal moves while planning in your head.
The main thing that's slightly "hidden under the rug" is that for "masking" to make sense in the first place, MuZero needs to know a set of all moves that may be legal at some point in the games.
I would love to see how this might go for more complicated games such as NES adventure games and RPGs.
This isn't just weak players. E.g. strong chess players often consider moves as if blocking pawns weren't there. They might consider a bishop to be on a strong diagonal despite there being a blocking pawn because they can imagine moves that would happen if that pawn would disappear.
It's no different than a monkey operating a chess computer that makes sure the monkey only performs legal moves. Your suggestion would be akin to suggesting that the chess computer would be affecting the monkey's mind so that it can only think in terms of legal chess moves.
It may be true, or perhaps it was a story concocted in order to emphasize his innate talent.
Unless they've changed a lot of things since the original paper, this is a bit exaggerated.
MuZero learns what moves are allowed in a given position/situation, but it still needs to know a finite overall set of possible actions.
E.g. for chess, it isn't told which fourty moves a available at each point in its search tree, but it still knows to only consider 64x64 discrete options.
It wouldn't know that moving its King from e1 to g1 must be accompanied by moving its rook from h1 to f1.
Or that moving a white pawn from e5 to d6 must in some cases be accompanied by removing the black pawn on e6.
I guess the environment does these in response. That doesn't suffice for
moving its pawn from b7 to b8 must be accompanied by replacing said pawn by some other piece.
https://deepmind.com/research/publications/Mastering-Atari-G... https://arxiv.org/pdf/1911.08265.pdf
If you're just messing around with 1 GPU and a desktop PC you should be happy to get Atari breakout to work.
They're just hyping up their Nature publication. Or did I miss something?
https://www.amazon.science/blog/neurips-reinforcement-learni...
NOT that one neural net can be trained to play all of the games.
So, while this is an interesting result and makes using the same architecture for specific applications easier and a bit more plug-and-play with little to no modification of the code, what it accomplishes is reducing the effort required by a software engineer or researcher on adapting the software before training even begins, but it pretty much still requires the same amount of training.
What this doesn't seem to do is allow the same trained network to be applied to multiple tasks (which I think most of the AGI comments are assuming), and it certainly doesn't generalize anything among the games it is trained on.
It would also give the gaming industry a kick in the pants to start making better AIs.
DeepMind made a good effort with AlphaStar at building an AI that could compete with top-level humans in Starcraft. It wasn't superhuman; it could still be consistently beaten by the absolute best Starcraft players, especially as Zerg or Terran. However, as Protoss, AlphaStar was truly a pro-level player. I'm somewhat surprised DeepMind didn't go further and try to optimize AlphaStar to truly be superhuman. I'm not sure if that indicates a fundamental limitation of their approach or whether it was a shift in approach. This was with successively refined limitations on AI action speeds that caused AlphaStar to really rely on strategy and tactics rather than brute force speed.
Regardless, real-time strategy games feel much more difficult than turn-based strategy games to develop a good AI for. Just being able to split things into discrete turns seems like a massive simplification.
I have a strong belief that it could grow and I’d like to contribute (and join the development)
It's an old discussion, and while a few of the deep learning results are really impressive I don't think any of them have fundamentally changed that discussion, yet.
You can easily handwave that all generated knowledge might be indirectly useful; i think that's fair but also different than the distinction I drew.
If you're curious how it would work, I implemented AlphaZero (but not Mu yet) using GBDTs instead of NNs here: https://github.com/cgreer/alpha-zero-boosted. Instead of saving the "weights" for a GBDT, you save the split points for the value/policy model trees, but the concept is the same.
You're welcome.
> Why did you use GBDTs instead of NNs?
I mostly wanted to build an implementation to see how it worked; I was more familiar with GBDTs than NNs, so I figured I'd start with that. At its heart, AlphaZero is the marriage of two great ideas: using a Monte Carlo Tree Search (MCTS) to efficiently look ahead and find good moves and using a powerful ML model (like a ResNet) as a bot's intuition about which positions are good to be in (value network) and which moves are good when you're in which positions (policy network). So if a GBDT is powerful enough for your use case, the "ML Model" component in the MCTS+ML Model AlphaZero setup should be able to be swapped out with it if you want.
But I was also curious if GBDTs would do almost as well as a NN, because GBDTs can be much more efficient w.r.t. cost/energy. At the time when AlphaZero came out, I think it cost >$10M to train a superhuman Go algo. Nowadays KatoGo [1] can do it for <$50K. The most expensive part of training is the self play. You basically have bots play millions of games against each other and learn from the results of those games. Getting value/policy predictions each move from the ML models is a majority of the computation during self play, so if you make that more efficient, you should be able to train a bot faster/cheaper.
Check out this HN thread if you're interested in more AlphaX shenanigans: https://news.ycombinator.com/item?id=23599278
All are _very_ strong. KataGo is ungodly strong, it beats pros.
Learning Go is about more than just playing against strong players, but it could help. The biggest difficulty is that the strong AIs aren't actually that good at playing handicap games, and they're also almost completely unable to explain to you why you should play one move over another.
You wouldn't know that from watching JBXKataBot playing on KGS with a typical 5-7 stone handicap.
Last I saw I remember katago playing up to maybe 4 stones pretty well but the games being poorer quality over that.
That's the problem. To learn, we need an AI that can be just a little bit stronger than humans, but at the same time we need an AI that makes natural moves, not an AI that makes great moves 90% of the time and clear blunders 10% of the time.
And playing go, the AI should be able to give handicap stones and play reasonable teaching moves.
I tried learning Go a little while back but hit a wall. Was thinking about trying this more gamified option:
There does exist research on how to make a human-like weak player: https://arxiv.org/abs/2006.01855
The basic idea is to look at weak human games and try to predict when a mistake will be made. But I don't know if there's any approach that can do that without access to a corpus of human errors.
For training you really want your good moves to be rewarded and your bad moves pointed out, but if the AI just plays up or down to match what you do instead, there's no signal getting back to you on how you're doing.
I think the strength or lack thereof of your opponent is actually much less important than the strength of the AI you use to review your games. After each game you should study the AI's advice and learn the moves it recommends.
That said, I think Deep Mind should go all in for solving practical real world problems.
I read qm. But does this actually useful for partial info. This is also another life situation where you never have full information.
I still wonder about the intelligence.