Minigo: An open-source implementation of the AlphaGo Zero algorithm
github.com
github.com
I am finding it quite useful. (I'm not the author, just a happy reader! :-))
cf. https://gogameguru.com/i/2016/03/deepmind-mastering-go.pdf, Extended data table 2
(I'm not much better :)
To get this working:
* Acquire a GTP-capable GUI, such as Sabaki
* Acquire the latest Leela Zero release
* Acquire a recent Leela Zero neural net
* Set up Sabaki to use LZ with the net passed as an argument, e.g. "-t 1 -p 1600 --noponder --gtp -w d16fa4c3801e55ec21e0df7ead67980fe8d4ee49188a3516818207ad28b017a6"
It's a bit of work but nothing too hard. I should mention that this may require a semi-decent GPU (my old GTX 750 works fine).
A nice win against Gnuchess (a very weak opponent, but nonetheless :) - https://github.com/glinscott/leela-chess/issues/47#issuecomm...
Not sure why you guys don't show the PGN, but here you are:
1. e4 Nc6 2. Nf3 Nf6 3. Nc3 d5 4. exd5 Nxd5 5. Bb5 Nf4 6. O-O Bf5 7. d4 Nd5 8. Ne5 Qd6 9. Nxd5 Qxd5 10. Bxc6+ bxc6 11. c4 Qd6 12. Qf3 g6 13. Nxc6 Bg7 14. Bf4 Qe6 15. Rfe1 Qxc4 16. Rxe7+ Kf8 17. Rae1 Kg8 18. b3 Qc2 19. Qd5 Be6 20. R7xe6 fxe6 21. Qxe6+ Kf8 22. Bh6 Bxh6 23. Qf6+ Kg8 24. Ne7#
Question : when you switch to self-play reinforcement learning, do you plan on starting from the networked obtained in supervised learning or tabula rasa? I understand starting from tabula rasa will require more comptuting power/time, but if you start from the supervised learning network, isn't there a risk you inherit human biases in the game style? It would also defeat the purpose of having the system discover existing chess theory and possibly new one.
Should be fun to watch it learn chess theory :).
> very carefully respects all possible game-ending pathways
Not sure exactly what this one means but it implements a is-game-over method, so you could do MCTS with it.
> a translation of a chess board into an array that a NN could understand
Chess positions are represented by sets of 64-bit integers so I don’t think this would be a blocker.
> and a schema by which to flatten the array of all possible moves (both legal and illegal) into a single vector
The list of legal moves is a set of bitboards as well.