How AlphaZero Mastered Its Games
newyorker.com
newyorker.com
It was almost a year ago that lc0 was launched, since then the community (led by Alexander Lyashuk, author of the current engine) has taken it to a totally different level. Follow along at http://lczero.org!
Gcp has also done an amazing job with Leela Zero, with a very active community on the Go side. http://zero.sjeng.org
Of course, DeepMind really did something amazing with AlphaZero. It’s hard to overstate how dominant minimax search has been in chess. For another approach (MCTS/NN) to even be competitive with 50+ years of research is amazing. And all that without any human knowledge!
Still, Stockfish keeps on improving - Stockfish 10 is significantly stronger than the version AlphaZero played in the paper (no fault of DeepMind; SF just improves quickly). We need a public exhibition match to setttle the score, ideally with some GM commentary :). To complete the links you can watch Stockfish improve here: http://tests.stockfishchess.org.
Of course I can't be sure, because Google refuses to give out anyone access to alphazero, or a network trained with it. Personally, that gives me more confidence they know there are significant exploitable weaknesses.
You can download the weights for LCZero right now though and try out your theory. https://github.com/LeelaChessZero/lc0/wiki/Getting-Started
I'd prefer to try with a go player, because as you say, in chess it's hard to exactly control the input to the network, it's easier in Go.
There's current best weights available. Not alphazero, but I would expect that issues would be general and so if there are issues with leela zero they may transfer and if you don't see issues with leela zero they're unlikely to exist in alpha zero (at least, if they do they may be very particular to subtle training differences).
Would be very interested to see what you find if you get the chance.
If you want to construct a particular position on the board, you'd likely need to use multiple steps, require the AI to play very particular moves and then the outcome would be a certain move from the AI. Even then, a simple incorrect classification doesn't help all that much, you need your opponent to make repeated mistakes.
I think in reality if you uncovered a type of move it wasn't expecting you are likely to uncover a new strategy in general rather than a trick. Image classification however lets you play uninterrupted with tiny pixel value changes, and you only need a single incorrect output to "win".
For alphazero, the input is the board, which you can't manipulate arbitrarily. You can run an evaluation of a board based on a move and see if its significantly different than the evaluation that alphazero comes up with, and maybe try to exploit that. But if you have a better evaluation of some state than that of alphazero, you're likely a stronger player anyway so this extra step is unnecessary. Most of the value of the bot comes from the evaluation function of a board, along with some hyper-parameters. But the evaluation is probably the most important part and the most difficult to replicate.
Generating the right noise was proven to be successful against NNs (https://blog.openai.com/adversarial-example-research/) but I am not sure how could you apply that to this context.
Malcolm Gladwell, as well as many other great writers, can be found here:
https://www.newyorker.com/contributors/malcolm-gladwell
This weeks story on Donald Trump might be of interest:
https://www.newyorker.com/magazine/2019/01/07/how-mark-burne...
I highly recommend the AlphaGo movie as well, it does a great job documenting the psychology of professional Go in the world of AI.
Instead of using gender neutral pronoun like "they", the author used a feminine pronoun.
The New Yorker style guide allows the author to use he or she at their discretion. That’s been the norm for hundreds of years.
The exact implementation details are probably kept secret, but the idea is to do a few steps of minimax / alpha-beta rather than completely random play in the playout phase of MCTS.
This makes me think that the contribution of AlphaZero is not necessarily neural nets, but rather MCTS as a succesful method to search the game tree efficiently.
[1] http://tcec.chessdom.com/ [2] http://www.chessdom.com/komodo-mcts-monte-carlo-tree-search-...
I can't find this claim in the linked paper. What I can find is a statement that AlphaZero has demonstrated that 'a general-purpose reinforcement learning algorithm can achieve, tabula rasa, superhuman performance across many challenging domains'.
Personally, and I'm sorry to be so very negative about this, but I don't even see the "many" domains. AlphaZero plays three games that are very similar to each other. Indeed, shoggi is a variant of chess. There are certainly two-person, zero-sum, perfect-information games with radically different boards and pieces to either Go, or chess and shoggi - say, the Royal Game of Ur [1], or Mancala [2], etc, not to mention stochastic games of perfect information, like backgrammon, or assymetric games like the hnefatafl games [3], and so on.
Most likely, AlphaZero can be trained to play many such games very powerfully, or at a superhuman level. The point however is that, currently, it hasn't. So no "demonstration" of general game-playing has taken place, and of course there is no such thing as some sort of theoretical analysis that would serve as proof, or indication, of such ability in any of the DeepMind papers.
I was hoping for less ra-ra cheerleading from the New Yorker, to be honest.
________________
[1] https://en.wikipedia.org/wiki/Royal_Game_of_Ur
However, with the framework they've built it's easy to verify the claim for new games (given sufficient computational power). Maybe that should have been the point made.
- asymmetric games, like hnefatafl, that probably can be covered — considering that AlphaZero can handle a late-stage situation with asymmetric options;
- what I understand to be stochastic games, of dice-based games, like the Royal Game of Ur, Mancala and backgrammon. I would expect that you have to re-define success by representing risk profile in the strategy; that could introduce complexity that the network can’t handle.
It's AlphaZero's deep neural net component, that's used to learn an evaluation function and move orderings that will need a substantial redesign to take into account imperfect information. The difficulty of this redesign will vary considerably between games- in some games, information is gained throughout the game, by observing another player's moves (e.g. in Poker), in some others an initial state (e.g. a starting deal in card games) dominates the probability that a certain possible board state is the real board state (e.g. Bridge) [1].
On top of that, AphaZero's deep net has the shape of the board and the legal movements of pieces on it hard-coded, as part of the net's structure. That would also need a substantial redesign to accommodate a card game, or any other kind of game without a board and without pieces that move on it. In fact, different card games will require different architectures, most likely. It's very hard to see how, e.g., the same neural net structure could be used to encode both Bridge and Poker rules - and still allow learning chess, shogi and Go.
Given the great variety of board games out there (and that's only classical games, I'm not even considering modern board games, like Settlers, etc) a lot of very hard work would be required to even train AlphaZero to play any game that's not very similar to chess, shogi and go. Not to mention, training AlphaZero is very expensive (wikipedia quotes a cost of $25 million for AlphaGoZero, AlphaZero's predeecssor, and that's just to buy the hardware [2]). So I don't see how or when they'll demonstrate the "general" game playing power of their system.
Basically, I think all that stuff about "generalized" game playing is just so much pointless bragging. The way DeepMind designed AlphaZero is exactly how everyone else has designed their systems- hard-coded with structures appropriate to the targeted game (e.g. boards and pieces, etc). DeepMind are clever in that they chose three very similar games, and then threw an immense amount of money on the problem of solving them all in tandem. And still they had to train different models for each game. That's just no way to get to general game playing.
___________
[1] See chapter 5. Adversarial Search in AI: A modern approach, 3d Ed. for a discussion of imperfect information and stochastic games and the difficulties of designing evaluation functions for them.
[2] https://en.wikipedia.org/wiki/AlphaGo_Zero#Hardware_cost
Honest question:
Are games like backgammon really considered "perfect information" in the sense relevant here?
No player has any secrets from the other, but neither knows what the dice will do, which certainly is important information.
Is it impossible to apply to these types of games? Every time I read about AlphaZero the articles mention that the techniques are meant for games of perfect information.
EDIT - The David Silver lecture I mentioned actually mentions a Scrabble AI, Maven, which successfully applies MCTS. Here's a link: http://www0.cs.ucl.ac.uk/staff/d.silver/web/Teaching_files/g...
https://deals.manning.com/go-comp/
It’s been really fun to work through the books.
> Before there could be acceptance, there was depression. “I want to apologize for being so powerless,” he said in a press conference.
Lee Sedol was clearly upset, especially after the first two matches, but I think that apology was more out of politeness than depression, really.
At a glance, the parameters of the match seem unfair to me -- and tilted heavily towards AlphaZero. If the code, were open source, this would not matter; anyone could run a rematch. As it is, I haven't seen any convincing evidence that AlphaZero is stronger than Stockfish when Stockfish is allowed to use its full breadth of knowledge and run on equal hardware.
Second paper they played fair even giving stockfish a 10 to 1 advantage on same hardware as tcec uses.
As for the GM side, Nakamura basically said the same.
Can I just state again, because of your aggressive tone in multiple comments now, that I do believe AlphaZero is stronger, I don't believe there are real shenanigans going on, but it's _still_ sad that we can't reliably, publicly verify this stuff.
This claim implies foul play, I asked for a shred of evidence. Call me aggresive if you want, I just can't stand this kind of bullshit.
This is all either of us would have done if we were peer reviewing the paper, don't really understand the hostility about trying to reproduce a published paper.
Matthew Sadler on chess24
https://www.youtube.com/watch?v=JacRX6cKIaY&list=PLAwlxGCJB4...
Daniel King
https://www.youtube.com/watch?v=pFtY7gNRVRI&index=3&list=PLh...
Anyway, I wouldn't be surprised if AlphaZero lines have existed at the top of the game for some time. Would be a no brainer for someone to have made Google an offer after the first paper.
I think there's a huge amount of cognitive dissonance going on, so that people can label AlphaZero's play more 'human'.
or:
if AlphaZero lines have existed at the top of the game for some time.
The second point, I'm saying it's likely some teams, possibly Caruana's have had access to AlphaZero already, and if the public had better access to its analysis we might be able to look at more games from 2018 and see its insights pop up. Certainly some of the lines Caruana played at the World Championship match were strongly backed by AlphaZero in the videos released, but weren't Stockfish's choice, for example.
If there's no public tournament, this might as well not have happened. I do not understand why Google is always special. Other engines are open, Google can test against Stockfish but not vice versa.
All these web companies take, take, take from Open Source and rarely give back.
I'll write a paper now that I beat Carlsen, but I'll refuse to do so in public.
Do you mean me? If so, why not say so. "Could people please stop X-ing here?" is very passive-aggressive. I looked up 'FUD' - Fear, uncertainty and doubt. Not sure how what I said counts as any of those. Your comment certainly seems to want to spread FUD, however. (And why is this your account's only comment on HN?)
I don't know the significance of your first sentence either. I know nothing about the versions of Stockfish used for this or anything else. Maybe you're taking too much for granted. i.e. that I know enough of the minutiae to understand your comments. Could you fill in the dots a bit? And who are 'all these web companies'?
I'm not super-interested in AlphaZero (or computer chess generally) - haven't read any of the papers, for example. But there's a lot of talk about the various Alphas in the online chess world, since before it was playing chess, and I found these videos very impressive. And it's ridiculous to say "This might as well not have happened". Maybe true for you, but not for the chess world, at all.
You can still run the PGN of the published games back and find some places where Stockfish 8 will analyse its own moves and find them blunders, so it _would_ be good for all this to happen out in the open, but I don't think there's any large scale deception going on here. I think it's beyond reasonable doubt now that AlphaZero is easily as strong as they are claiming.
The "Patronizing 'Her'"
Almost invariably, when the author decides to use the patronizing 'her' instead of the gender-neutral 'they' it's written by a man.
I have noticed that people who exclusively use "his" rather than "her" or "their" tend to be men though.