Acquisition of chess knowledge in AlphaZero
pnas.org
pnas.org
> [...] However, sharing the AlphaZero algorithm code, network weights, or generated representation data would be technically infeasible at present.
Very interesting paper overall. However, the excuse that code sharing is "technically infeasible" is wearing thin nearly 5 years after the initial AlphaZero paper was released.
The representations are probably not manifest in a way that would be intelligible if shared.
I don't have an explanation for why they wouldn't share the weights.
It is possible that the entire codebase depends on Google-only infrastructure.
It's possible other infra is somewhat hairy to decouple, for example the code they use to allocate and use GPU resources is internal.
(I work on ML at Google and we use some of Deepmind's stuff)
> Many Human Concepts Can Be Found in the AlphaZero Network.
> We demonstrate that the AlphaZero network’s learned representation of the chess board can be used to reconstruct, at least in part, many human chess concepts. We adopt the approach of using concept activation vectors (6) by training sparse linear probes for a wide range of concepts, ranging from components of the evaluation function of Stockfish (9), a state-of-the-art chess engine, to concepts that describe specific board patterns.
> A Detailed Picture of Knowledge Acquisition during Training.
> We use a simple concept probing methodology to measure the emergence of relevant information over the course of training and at every layer in the network. This allows us to produce what we refer to as what–when–where plots, which detail what concept is learned, when in training time it is learned, and where in the network it is computed. What–when–where plots are plots of concept regression accuracy across training time and network depth. We provide a detailed analysis for the special case of concepts related to material evaluation, which are central to chess play.
> Comparison with Historical Human Play.
> We compare the evolution of AlphaZero play and human play by comparing AlphaZero training with human history and across multiple training runs, respectively. Our analysis shows that despite some similarities, AlphaZero does not precisely recapitulate human history. Not only does the machine initially try different openings from humans, it plays a greater diversity of moves as well. We also present a qualitative assessment of differences in play style over the course of training.
Something like GPT-3 can do multi-digit arithmetic much better than chance, giving results for values it was certainly never trained on. Similarly, transfer learning, where you start training a model on some input less related to the task, and then switch to inputs closer to your task at the end, can substantially reduce total training time. The task can be radically different; to use GPT-3 as an example again, compared to starting with a completely randomized model, it reduces training by a factor of about 10x to go from PCM audio samples encoded as text patterns, or abstract art bitmaps encoded as text patterns, to English text. GPT-3 is learning something about arithmetic. It's learning something that is common to music, abstract art, and English text. It might be as simple as basic patterns from geometry and arithmetic (that's my guess). But no one could even begin to point you in the direction of what that structure it is teasing out really is.
I remember with Deepminds breakout AI one very easy way to see the difference to human play was to change the shape of the paddle. Even very slight changes completely threw the AI off, so it was obvious it hadn't understood the 'breakout ontology' in a human way.
I'd expect the same from chess. Humans who understand chess at a high level well obviously play worse in non-standard variants but the familiar concepts are still in play. If an AI has a human-like grasp of high level concepts it ought to be pretty robust to some changes to the game rules like changing the dimensionality of the board.
In fact I'd argue it's unfair to expect an AI to automatically generalize without attempting to train it for that.
The only reason humans can quickly learn many games is that they've been exposed to a variety of tasks throughout their lives, as well as the fact that we have innate biases that we might use when designing games by humans for humans.
AlphaZero is also more interesting when you want to know how a general game-playing AI trained from scratch approaches the game.
I think using pre-NNUE Stockfish is partly because the classic Stockfish evaluation function has a lot of human knowledge explicitly built in in already interpretable ways, making it a good contrast for comparison.
https://en.chessbase.com/post/acquisition-of-chess-knowledge... https://arxiv.org/pdf/2111.09259.pdf
The hypothesis is some children will deeply embed the algorithms into their own playing style - leveraging the subconscious to the greatest degree possible. Basically, we are training the human mind in the same way that we train AI. Would it work? Probably not, but our current approach (studying openings, etc.) is obviously not working so it makes sense to try something new.
"Probably not" indeed. Human brains need to be stimulated in order for them to build their neural net. Blindly following instruction is not stimulating.
>but our current approach (studying openings, etc.) is obviously not working
That's part of the current approach, not the full approach. Kids that show promise in Chess, already train with computers today (in addition to personalized coaching, and strategy study). I have no doubt that the current generation of chess players is best of all time, and that the next generation will be even better. Even with that, I don't think any human will be able to train enough to beat even Stockfish, much less Alpha Zero - just as no human will ever train enough to beat computers at arithmetic.
As for how close Stockfish 8 is to a current version of Stockfish, Stockfish 15 has ~400 more ELO than Stockfish 8 as measured in their own self-tests[4].
However, when making these comparisons between AlphaZero and Stockfish it might be tempting to frame it as machine-learning-engine vs classical-engine, which is not true. Stockfish incorporates a neural network when evaluating chess positions[5].
[1]: https://en.wikipedia.org/wiki/Top_Chess_Engine_Championship#...
[2]: https://en.wikipedia.org/wiki/Stockfish_(chess)#Top_Chess_En...
[3]: https://en.wikipedia.org/wiki/Leela_Chess_Zero
[4]: https://github.com/glinscott/fishtest/wiki/Regression-Tests#...
[5]: https://stockfishchess.org/blog/2020/introducing-nnue-evalua...
[1] https://www.youtube.com/watch?v=N2MncpRMnFA&feature=youtu.be...
Think of the amount of complex computation must be happening for an elite gymnast for example. Sure, they aren't solving the equations consciously but at some level computation is happening - or even something as simple as making sense of all of the individual photons arriving at the eye. The subconscious is what I am suggesting we try to exploit. Even then, it seems likely that a "special" brain would be needed - hence why large numbers would be required in order to find it.
By that logic, my cat should be trainable to a Chess Grandmaster level because she performs complex computation as she navigates my backyard.
>The subconscious is what I am suggesting we try to exploit.
Is there any evidence that a subconscious is exploitable it that way ... at all. Because learning doesn't work that way - especially for higher-order knowledge like Chess.
The biggest failing would be, of course, the children would not know how to play against moves that are not optimal.
This is the most important part of chess knowledge: given some suboptimal play, prove it is suboptimal by defeating the player who made the suboptimal move.
Otherwise, as some chess schools do, they would be repeating famous openings by rote, without understanding the meaning behind each move in the opening.
This ability would matter little when playing against Alpha Zero, but it would make all the difference when playing against humans.
https://www.reddit.com/r/chess/comments/c4dgas/alpha_zero_fi...
I don't think they mention anything about speeding up training. The original training time for AlphaZero was only a few hours anyway so I don't think that was ever a major constraint. I would imagine that each neural network performed best on the variant on which it was trained.