Show HN: ChessCoach – A neural chess engine that comments on each player's moves
chrisbutner.github.io
chrisbutner.github.io
It's a chess engine with a primary neural network just like AlphaZero or Leela Chess Zero's, but it adds on a secondary "commentary decoder" network based on Transformer architecture to comment on positions and moves. All of the code and data for training and search is from scratch, although it does use Stockfish code to generate legal moves and manage chess positions.
You can watch it play on Lichess here: https://lichess.org/@/PlayChessCoach/tv or challenge it here: https://lichess.org/?user=PlayChessCoach#friend, and see its commentary in spectator chat. It only plays one game at a time, so you may need to wait a little bit. It's fairly strong (~3450 rating, roughly on par with Stockfish 12 or SlowChess Blitz 2.7), but you can set up a position when challenging it so that it's missing a couple pawns or a piece (Variant: From Position).
I ended up writing much more about it than I expected. If you're into the technical side of chess or machine learning, beyond the linked overview, there's:
High-level explanation: https://chrisbutner.github.io/ChessCoach/high-level-explanat...
Technical explanation: https://chrisbutner.github.io/ChessCoach/technical-explanati... (including code pointers)
Development process: https://chrisbutner.github.io/ChessCoach/development-process... (including timelines, bugs and failures)
Data: https://chrisbutner.github.io/ChessCoach/data.html (including raw measurements and tournament PGN files)
And the code is here: https://github.com/chrisbutner/ChessCoach (C++ and Python, GPLv3 or later)
Happy to answer any questions!
I'm actually hopeful that some search techniques such as SBLE-PUCT[1] or better derivations can make their way into other open source projects, but they've had big teams working for a while on similar, often better ideas, so we'll have to see.
[1] https://chrisbutner.github.io/ChessCoach/high-level-explanat...
Partly because of the way it tries to search more widely to avoid tactical traps, it can also be a little sloppy in holding advantages or minimizing losses (this could use some more work and tuning). This ends up making it a little drawish, so it loses less than you'd expect to Stockfish 14, but also doesn't beat up weaker engines as well as Stockfish 14 does.
You can see some of this in the raw tournament results[1]. At 40 moves per 15 minutes, repeating, each engine draws with the ones above and below it, but starts to win and lose at a distance of 2 or 3.
At 5+3 time control, ChessCoach goes 1-0-29 vs. Stockfish 12, but Stockfish 12 is better at beating Stockfish 8-11 than ChessCoach is, so CC ends up between SF11 and SF12 in the end.
On Lichess, where there's no "free time" to get ready for searches, ChessCoach's naïve node allocation/deallocation makes it waste time, and means it can't ponder for very long on the opponent's time - a big opportunity for improvement (it needs a multi-threaded pool deallocator that can feed nodes back to local pools for the long-lived search threads). I think it's also hitting a bug with Syzygy memory mapping that Stockfish works around via reloading every "ucinewgame" (which I don't trigger on Lichess). So, overall, its performance on Lichess is worse.
Also, you can't read too much into this data - very few games, and no opening book.
[1] https://chrisbutner.github.io/ChessCoach/data.html#appendix-...
Why this limitation? Is it fairly computationally expensive to run?
Lots of room for improvement!
Next steps could be using one of Lc0's backends for GPU scenarios, or taking the other side of the trade and using the C++ API for TPU.
There's also your typical CPU and memory optimizations that could be made - some baseline work there but not targeted.
Getting deep into RL specifically wasn't so necessary for me because I was just replicating AlphaZero there, although reading papers on other neural architectures, training methods, etc. helped with other experimentation.
You may be well past this, but my biggest general recommendation is the book, "Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow" to quickly cover a broad range of statistics, APIs, etc., at the right level of practicality before going further into different areas (for PyTorch, I'm not sure what’s best).
Similarly, I was familiar with the calculus underpinnings but did appreciate Andrew Ng's courses for digging into backpropagation etc., especially when covering batching.
To me it's probably OK to train a model on this, at least for hobby purposes, though some GitHub Copilot critics might disagree. And a large part of ChessBase's business model is based on ripping off other people's IP and presenting it as their own [1]. But still, I can see why the author might want to be coy about answering this question.
[0] https://en.chessbase.com/post/new-mega-database-2021
[1] https://lichess.org/blog/YCvy7xMAACIA8007/fat-fritz-2-is-a-r...
I had a nice long conversation with two of the authors of [0] at ACL.
One thing we discussed was the reverse problem. That is, as a player, could I give commands to the model and have the engine figure the moves that would best satisfy them.
This ranges from concrete like "take the black square bishop" (there is still variability like which piece should take it or if it's even possible) to more complex positional stuff like "set up to attack the kingside."
Any thoughts on this line of research?
[0] Automated Chess Commentator Powered by Neural Chess Engine (Zang, Yu & Wan, 2019) https://arxiv.org/pdf/1909.10413.pdf
I think this line of thinking could eventually lead to automated metrics for commentary evaluation, which could in turn lead to better methods than top-k/top-p for turning a bunch of sequential logits into a sentence or paragraph - basically treat it like MCTS/PUCT also.
The problem is that if you look at high-level commentary - maybe Radjabov-MVL on https://www.chess.com/news/view/2021-champions-chess-tour-fi... (I'm not the best judge, just a quick search) - it's not often possible to predict the move starting with the comment. And if you did, you might end up with very dry metrics and reverse commentary.
But this direction has a lot of potential I think, beyond just chess, into more of an algorithmic/generational support for pure NN-based language models.
how much does a game cost in CPU time money?
How do I get the commentary for a game I played? Oh, it's in Analysis page.
It plays chess very well, but the commentary is incoherent and doesn't match the game well -- The attacks described are nonsense and the coordinates are wrong. It seems a little confused about which side is which? It thinks a rook can diagonally attack a bishop, and seems to name squares opposite from their actual name.
Costs are difficult to work out - it depends on cloud vs. self-hosting, what kind of TPUs/GPUs, how long you're calculating over.
The advantage that classical/NNUE engines have is that they can more easily spread over distributed frameworks like Fishtest.
Agreed, this looks superficially like commentary on the game, but honestly it doesn't seem more pertinent to the game score than a Markov chain trained on all the commentary would be (presumably this isn't true, and the author started with something like that Markov chain and the current version is way better in terms of some fitness function).
I wonder if there just is not enough training data available. GPT-3 overcomes this by harvesting a ridiculous amount of training data. AlphaZero, and the chess engine here, which is excellent, overcome it by generating their own training data through self play. But that's not applicable to the task of generating commentary.
Installation for GPU is covered here: https://github.com/chrisbutner/ChessCoach#installation (a little messy, sorry)
It seems quite good overall, but occasionally there are fun AI-isms like this :)
e.g. on move 22
I don't know why black played this. As black, I would've played 27...Kf8 to protect f6 and e5.
the 27..Kf8 is a link ... but the game hasn't proceeded to move 27 yet, so the link goes nowhere.
Clearly it's trying to show some line that ends with 27....kf8 ?
Anyway, the commentary itself is pretty excellent, great job!
And thank you!
I don't think this comment was accurate:
> 23. c4: "White tries to get his pawns moving. I am still thinking that I have to move my bishop, but that is too slow. I think white should have moved his king back to c3 to prevent my pawn from becoming a passer."
The idea behind moving the pawn was because black playing c4 would have instantly lost both pawns since the bishop has nowhere safe to go where it still defends c2 (Bh7 leading to g6 losing the piece.) I don't think the king could have made it to the c file in time to stop that.
> 34... Rxg4: "Now I can't stop him from advancing his pawn."
Is this meant to be speaking from whites perspective (since black just got the unopposed h file pawn.)
Interestingly only 10. O-O links with a move, I think it would be really helpful if they all linked to moves. Also i'm really excited for this kind of analysis! It would be really cool if request a computer analysis eventually generated such analyses trying to figure out each sides ideas.
I wonder if we are getting snippets of variations in some cases.
There's real substance here. Well done. I hope you keep developing it, you're on to something novel.
That's overall less general than "feed position network output into a transformer", but presumably less data-constrained.
And if you spice up the commentary sampling parameters, it gets even more inventive, making up names, and saying that "the rook is pinning Fischer against the king".
So you're telling me ... There's a chance?
[1] https://www.chess.com/news/view/updated-alphazero-crushes-st...
> It plays chess with a rating of approximately 3450 Elo... [compared to] Stockfish 14 at 3550 Elo.
So for a non-technical audience, I feel like it's easier to give a ballpark that they can understand without having to pull in too much context around Stockfish, CCRL, etc. It may have been better to clarify further in the docs though.
The "Data" document does give the relative Elo breakdown in the appendices.
A remark about opening preparation: The best metaphor I've seen here is the one about snooker. Ronnie O'Sullivan needs a good safety game because his opponent can clear the table. You don't.
> This time control is too fast for me, please challenge again with a slower game.
Is that expected?
You can challenge https://lichess.org/?user=chesscoachclassical#friend to 30+20.
Unfortunately, neither of them support correspondence.
One bug found: https://lichess.org/NvbQTf2O/black
> PlayChessCoach 23. Nc7: "The point of White's previous move. The knight on d7 is trapped and the rook on a8 is not protected."
There is no knight on d7 or no knight that could be moved there. (Or so it seems to me.)
In that same game this was funny banter:
> PlayChessCoach careful not to get mated in the long run."
Thank you for writing architectural overview, enjoying it!
That is, in addition to the usual evaluation and policy heads, this takes the intermediate board representation and outputs a seed vector that is fed into a transformer text generator?
Or do other things go into the seed? Like the search tree somehow? Otherwise I suppose the commentary will not be able to comment on deeper tactics?
Or maybe this doesn't work using a seed vector at all, but with a custom integration from the board into the transformer somehow?
So, now the commentary decoder is just trained separately on the final primary model. The previous and current game positions are fed into the primary model, and the outputs are taken from the final convolutional layer, just before the value and policy heads. Then, that data plus the side to play is positionally encoded and fed into a transformer decoder.
It would be better for a search tree/algorithm to be used for commentary too so that tactics could be better understood, but that would need some kind of subjective BLEU equivalent, and metrics like those don't work well for chess commentary.
You can see a diagram of the architecture here: https://chrisbutner.github.io/ChessCoach/high-level-explanat...
Actually, I can't figure out from your explanation why you trained the whole network yourself instead of just using Leela's network and training the commentary head on top?
If you wanted to in-cooperate the search, maybe you could just take the 1800 or so probabilities output by the MCTS and add some layers on top of that before concatenating with the other data fed into the transformer.
In either case, this is a fantastic project and perhaps an even more impressive write up! Congrats and thank you!
In terms of search-into-commentary, concatenating like that may be interesting, as long as it can learn to map across - definitely plausible without too much work. I was originally thinking of something more complicated, combining multiple raw network outputs across the tree through some kind of trained weighting, or additional model via recurrence, and punted it.
Ignore my BLEU comment, mixed those up between replies - that was the other potential use of search trees for commentary, an MCTS/PUCT-style alternative to traditional sequential top-k/top-p sampling, once you have logits and are deciding which paragraph to generate.
Thanks!
I'm particularly interested in the second neural net that generates explanations. Is it only a neural net to generate the natural language using artifacts from the original engine net, or is it actually inspecting the state of the engine net to derive the insights?
It's an interesting application of explainability of AI algorithms that Christian talks about in his chapter on transparency. In particular, he discusses "saliency" of algorithms (knowing what parts of the input were most important in producing a prediction) and "multitask nets" that output multiple predictions (so, maybe here, one output is the best move, and another output is the explanation).
The writeups are fantastic reading. I see almost no sources in common with the bibliography in The Alignment Problem (which has a 50 page bib) which makes them nice complements. Only common citations I could find were Sutton 1988 (Temporal Differences) and Silver et al 2016 (AlphaGo), 2017 (AlphaGo Zero) and 2018 (AlphaZero).
There's a note (44) from chapter 5 entitled "Shaping" where Christian talks about "meta-reasoning: the right way to think about thinking. When you play a game–for instance, chess–you win because of the moves you chose, but it was the thoughts you had that enabled you to choose those moves...Figuring out how an aspiring chess player–or any kind of agent–should learn about its thought process seemed like a more important but also dramatically harder task than simply learning how to pick good moves." This note might as well be direct inspiration for this project. It goes on to quote Stuart Russell about "a computation that changes your mind about what is a good move to make...reward that computation by how much you changed your mind...so you could change your mind in the sense of discovering that what was the second-best move is actually even better than what was the best move." That's in the context of a cautionary tale where only optimizing for those "changes of mind" doesn't necessarily find a correct outcome, and that you have to "arrange these internal pseudorewards so that along a path, they add up to the same as true, eventually." This sounds pretty much like the task of a coach.
Deeper introspection is a really important goal, but by the time you make serious progress there, chess is the least of your worries.
I do really like the work people have put into introspection and visualization so far though: DeepDream comes to mind. There was also another great paper or page that I can't find.
With no bishops on the board...
"This is a good move for black. It attacks the center and attacks the pawn on e4. It also allows the development of the knight to c6."
"The main line. I don't know why I played this. But, I know that it's not a good idea to block in your own pieces, and that's what I would've done. But, this is a mistake because of what I'm about to do." (This is after its own move, and sounds very GPT3-ish.)
"He decides to kick my knight, but this is not a good move. It is also a very good move because it threatens my knight and threatens my pawn on e5. I am not worried about the knight fork on f2."
"I'm not sure what I'm doing, so I thought that I could castle, and try to defend it with a pawn, but I don't think it's worth it."
"The pawn on e5 is pinned to the king, so it is time to castle."
These samples are from the first ten lines or so of a single game.[1] They're all like that. Don't know what you were reading.
Unfortunately it's very difficult to track down training data for chess commentary in the first place, let alone trim down biases. For reference, I was able to gather about 1 million samples, but it really needs a billion.
Hopefully through data augmentation and better general intelligence models we can make better progress on bias issues soon, as that's a huge problem when we start trusting AI models too much in life.
I think for the most part, it knows more than it lets on, but finding the right sampling methods (or better yet, generalized search) to generate the best comments is a tough problem because it's difficult to evaluate quality.
There's some info on the sampling methods here: https://chrisbutner.github.io/ChessCoach/high-level-explanat...
Harder would be more general models like GPT-2 and GPT-3.
Realistically, it's really hard to fix this encoded bias in language models.