DeepMind's MuZero teaches itself how to win at Atari, chess, shogi, and Go
venturebeat.com
venturebeat.com
If I'm interested in building a toy version following the Deepmind spec, which can be trained to reach super-human capabilities on a particular board game (Reversi, Chess, checkers, possibly even Go given enough compute), which of these "versions" of the project would be the easiest for me to understand/implement? (assume I have a basic understanding of the high-level concepts and lots of enthusiasm, but I'm not an expert).
My understanding is, AlphaZero is not just stronger than AlphaGo, but architecturally simpler and more efficient. That's what I'm looking for -- the implementation with the highest result/difficulty ratio.
Recently, Katago has reached similar levels of strength using a small fraction of the resources: https://arxiv.org/abs/1902.10565
It depends on what you mean by "more efficient." The significance of AlphaZero was that you can reach good results in a variety of domains even without human expert knowledge to provide supervised learning data or engineer features. It's efficient in terms of engineering resources.
A precisely tailored approach can always get better results.
I have the same question. Not sure I have an answer yet, but this paper includes some pseudocode that implements the algorithm: https://arxiv.org/src/1911.08265v1/anc/pseudocode.py
I'm planning on trying to train something simple like TicTacToe to both see if it works and understand how it works.
These one off experiments make it hard to know if AI is truly progressing or not. Naively I'd assume due to the decision tree leaves growing exponentially with depth, then we are facing an inherently unscalable problem, and we only are getting current gains due to advances in hardware, but the gains are only linear with exponential hardware improvements, especially if Moore's law is giving out and even with parallel computation we might end up turning the earth into a giant GPU array before we can reach parity with human intelligence.
Because of this I am presuming.
> In Go, MuZero slightly exceeded the performance of AlphaZero, despite using less computation per node in the search tree (16 residual blocks per evaluation in MuZero compared to 20 blocks in AlphaZero). This suggests that MuZero may be caching its computation in the search tree and using each additional application of the dynamics model to gain a deeper understanding of the position.
It also strikes me as possible that just not giving the system the rules to start with might have allowed it to explore more efficient strategies.
But applying that in the real world takes years.
It is even possible that our brains just do the same thing BTW. How many times do you run that scenario of a job interview in your head before you go there? How many times does it run in your subconscious virtually? How many times does it happen in dreams? And more profoundly, how often do those scenarios in our head are very inaccurate and simplified, and yet they still help us act in the real world nonetheless?