Human-like neural network chess engine trained on lichess games
maiachess.com
maiachess.com
1. It brings out the queen early attacking the f7 pawn when black plays sicilian.
2. So far has gone king's indian almost every time against d4 (e.g. catalan) and then has failed to challenge the center, instead going for a king side attack (exchanging bishop for knight to try to open the B file.
3. It sometimes drops its queen in complex exchanges and hidden attacks.
4. It sometimes drops pieces in complex exchanges (can't count how many pieces are covering a square).
5. It will exploit mistakes that *I can see*, unlike stockfish level 7+ where it will take you down an insane convoluted path to get you into such a bind to destroy your position (I've played GMs who ?can't?/don't do in this blitz).
6. Its attacks are shallow and lack depth, easily defended.
7. It sometimes moves a single piece too many times when it should be trying to advance its positional game (e.g. queen/knight), basically has very little positional play.
8. It sometimes pushes pawns aggresively to its detriment.
9. Won't resign in losing positions (*lol*).
a. Will play to the bitter end to try to get you to stalemate.
Couple things that could make it more human: 1. Moves fast! I think it actually moves too fast, there needs to be a better delay factor added in depending on rating. If we could make it think for a longer than usual amount of time after it finishes development, in complex positions, or when it's close to mate or about to lose an exchange. Also make it speed up when its getting low on time.
2. Make it randomly rage quit in a losing position like an asshole instead of resigning so you have to wait for the quit/disconnect detection and then claim victory/draw countdown (I jest, I jest, but if we do this, please make it some sort of setting).I even saw an IM vs. NM bullet game the other day where the NM was in a losing position but stayed in to grab a stalemate: https://www.reddit.com/r/chess/comments/kwoikt/im_not_a_gm_l.... Not sure if Levy was being unsportsmanlike to stay in the game despite being in a losing position, but even at a high level I think it's normal to play to the end if your opponent is in time trouble.
This kind of program seems like it would be much more satisfying to play just for fun, and perhaps (with a bit more analysis support) better still as a coaching tool.
A particular use case that's implied by the features is the ability to analyze errors that you would make as opposed to the exact errors that you made; as the personalized "Maia-transfer" model seems to have an ability predict the specific blunders that the targeted player is likely to make, those scenarios can be automatically generated (by having Maia play against Stockfish many times) and presented as personalized training exercises to improve the specific weak spots that you have.
38. Kxa5 Nxg2 39. Kb6 f5 40. h4 f4 41. h5 gxh5 42. Kc7 f3 43. Kd6 f2 44. Ke7 f1=Q 45. Ke8 Qe1+ 46. Kd7 h4 47. Kd6 h3 48. Kd5 h2 49. Kd4 h1=Q 50. Kd5 h5 51. Kd6 h4 52. Kd7 h3 53. Kd8 h2 54. Kd7 Qhg1 55. Kd6 h1=Q 56. Kd5 Ne3+ 57. Kd6 Nf5+ 58. Kd7 Ng7 59. Kc7 Qd1 60. Kc8 Qc1+ 61. Kd7 Qgd1+ 62. Ke7 Qhe1+ 63. Kf6 Nh5+ 64. Kg6 Nf4+ 65. Kf5 Nh3 66. Kf6 Nf2 67. Kg6 Kf8 68. Kf6 Ke8 69. Kg6 Kd8 70. Kg7 Kc8 71. Kg8 Kb8 72. Kg7 Ka8 73. Kg8 Ka7 74. Kg7 Ka6 75. Kg8 Ka5 76. Kg7 Ka4 77. Kg6 Kb3 78. Kg7 Ka2 79. Kg8 Ka1 80. Kg7 Ka2 81. Kg6 Ka1 82. Kg7 Ka2 { The game is a draw. } 1/2-1/2
These were bullet games where it was rated at 1700 and I am rated 1300ish...however I won a number of games against it. I never felt like I never had a chance.
I guess that part of the position space was undersampled in the training data!
One interesting thing to see would be how low-rated humans make different mistakes than Leela does with an early training set. How closely are we modeling how humans learn to play Chess with Leela?
Another thought: Leela, against weaker computers, draws a lot more than Stockfish. While Leela beats Stockfish in head to head competitions, in round robins, Stockfish wins against weaker computer programs more than Leela does.
I believe this is because Stockfish will play very aggressively to try and create a weakness in game against a lower rated computer, while Leela will “see” that trying to create that weakness will weaken Leela’s own position. The trick to winning Chess is not to make the “perfect” move for a given position, but to play the move that is most likely to make one’s opponent make a mistake and weaken their position.
Now, if Maia were trained against Stockfish moves instead of human moves, I wonder if we could make a training set that results in play a little less passive than Leela’s play.
(I’m also curious how Maia at various rating levels would defend as Black against the Compromised defense of the Evans Gambit — that’s 1. e4 e5 2. Nf3 Nc6 3. Bc4 Bc5 4. b4 Bxb4 5. c3 Ba5 6. d6 exd4 7. O-O dxc3 — where Black has three pawns and white has a very strong, probably winning, attack. It’s a weak opening for Black, who shouldn’t be so greedy, but I’m studying right now how it’s played to see how White wins with a strong attack on Black’s king. I’m right now downloading maia1 — Maia at 1100 — games from openingtree.com.)
1. e4 e5 2. Nf3 Nc6 3. Bc4 Bc5 4. b4 Bxb4 5. c3 Ba5 6. Ba3 d6 7. d4 exd4 8. O-O dxc3 9. Qd3 Nf6 10. Nxc3 O-O 11. Rad1 Bg4 12. h3 Bxf3 13. Qxf3 Ne5 14. Qe2 Bxc3 15. Bc1 Nxc4 16. Qxc4 Be5 17. f4 d5 18. exd5 Bd6 19. f5 Re8 20. Bg5 h6 21. Bh4 g5 22. fxg6 fxg6 23. Rxf6 g5 24. Rg6+ Kh7 25. Qd3 gxh4 26. Rxd6+ Kg7 27. Qg6+ Kf8 28. Rf1+ Ke7 29. Rf7# 1-0
This is their most recent ongoing head-to-head: https://www.chess.com/events/2021-tcec-20-superfinal
Current result: 9 draws, one win with Stockfish as White, and one win with Leela as White. Drawn.
There is also this one from a couple of years ago: https://www.chess.com/news/view/computer-chess-championship-... “Lc0 defeated Stockfish in their head-to-head match, four wins to three”. Stockfish did got more wins against the other computers, so won the round-robin, but in head-to-head games Leela was ahead of Stockfish.
stockfish finished +9 over 100 games.
I don't have any stock in those 2 engines, so I don't care which one is better than the other. At the end as a poor chess player it won't change anything :) It's actually interesting to compare how those two software are evolving and how they got here.
Stockfish is much older. And it took it a lot of hand tuning to reach its current level. It is (or was) full of carefully tested heuristic to give a direction to the computation. It would be very difficult to build an engine like stockfish in a short span.
Leela got there very very quickly. Even if it was not able to win in October, the fact that it got competitive and forced the field to adopt drastic changes in such a short period of time is impressive. It seems to be a good example of how sometimes no using the "best" solution could still be a win. Getting good results after a few months against something that required 10 years of work.
Right now, Stockfish is winning in the current TETC, but only by one point (one more win than Leela). https://www.chess.com/events/2021-tcec-20-superfinal Stockfish 12 27.5 - Leela 26.5
So it's an interesting result to me.
This actually may be the reason of higher ranking. What are the odds that a low ranked player will blunder a piece in a particular position? Quite low. But what are the odds that a low ranked player will blunder a piece in a game? They are quite high. So while this engine may predict most likely move, it can’t fake a likely game because it is too consistent.
I think GANs can be helpful to do something like this. One NN tries to make a man-like move and another one tries to guess whether the move was made by a human or the engine given the history of moves in a game.
I'm the main developer for Maia chess
Ideally you just use the sample as the basis, and then let an AI engine play against itself for training, and/or participate in real world games, such as they did with AlphaGo and/or AlphaStar.
If you sample from the probability distribution you are modeling, there is no reason it shouldn't play like a 1100 player.
I.e. most of the time if you leave your queen hanging and under threat your opponent will take it, but sometimes they just don’t see it. That’s the difference between playing a bot and a human a lot of the time - humans can get away with a serious blunder more often at low level play.
What part of that is unrealistic? This happens constantly to human players, and not only at 1100...
> Note also, the models are also stronger than the rating they are trained on since they make the average move of a player at that rating.
Reference: https://github.com/CSSLab/maia-chess
Let’s say you accidentally leave a pawn hanging and 90% of 1100 players would spot it, and 10% of the time they miss it. In this engine, as the 90% is the most likely move it spots it 100% of the time.
So this example shows that if you pick the most likely move for a 1100 player every move, you end up scoring better than a 1100 player. Of course it works the other way too on spotting brilliant moves that others miss, but I guess at 1100 level play there are more opportunities to mess up!
Detecting deepfakes and generating them are just adversarial training that will make deepfakes even better and then our society won’t trust any video or audio without cryptographically signed watermarks.
As an example, let's say there's a position where the best technical move will lead to a tiny edge with perfect play. Current programs like Stockfish and Alpha Zero will recommend that move. It would be better to instead recommend the move with a strong attack that will lead to a large advantage 95% of the time, even if it will lead to no advantage with perfect play. It seems one could extend Maia Chess to develop such a program.
If not, I wonder if that would make accuracy even higher!
My guess is no, because you have to get an exact output of a function which is not continuous at all. But maybe I am missing something?
It's great to see research on this.