Up to 10 or 15 moves, sure, we're well within common openings that could be regurgitated. By the time we're at move 20+, and especially 30+ and 40+, these are completely unique positions that haven't ever been reached before. I'd expect many more illegal moves just based on predicting sequences, though it's also possible I got "lucky" in my one game against ChatGPT and that it typically makes more errors than that.
Of course, all positions have _some_ structural similarity or patterns compared to past positions, otherwise how would an LLM ever learn them? The nature of ChatGPT's understanding has to be different from the nature of a human's understanding, but that's more of a philosophical or semantic distinction. To me, it's still fascinating that by "just" learning from millions of PGNs, ChatGPT builds up a model of chess rules and strategy that's good enough to play at a club level.
After reviewing the chat history I actually have to issue a correction here, because there were two moves where ChatGPT played illegally:
1. ChatGPT tried to play 32. ... Nc5, despite there being a pawn on c5
2. ChatGPT tried to play 42. ... Kxe6, despite my king being on d5
It corrected itself after I questioned whether the previous move was legal.
I was pretty floored that it managed to play a coherent game at all, so evidently I forgot about the few missteps it made. Much like ChatGPT itself, it turns out I'm not an entirely reliable narrator!
Qxd7 early on was puzzling but has been played in a handful of master games and it played a consistent setup after that with b5 Bb7. Which I imagine was also done in those master games. But interesting that it went for a sideline like that.
It played remarkably well although a bit lacking in plan. Then cratered in the endgame.
Bxd5 was strategically absurd. fxg4 is tactically absurd. Interestingly they both follow the pattern: Piece goes to square -> takes on that square.
This is of course an extremely common pattern, so again tentatively pointing towards predicting likely sequences of moves.
Ke7 was also a mistake but a somewhat unusual tactic with Re2 and f5 is forced but after en passant the knight is pinned. This tactic does appear in some e4 e5 openings though. But then the rook is on e1 and the king never moved or if it did, usually to e8, not e7. Possibly suggesting that it has blind spots for tactics when they don't appear on the usual squares?
Fascinating stuff.
But the presence of illegal moves doesn't really show that in my eyes. I fully understand the rules of chess, but I still occasionally make illegal moves. In 2017 Magnus Carlsen made one in a tournament [1]. The number of illegal moves suggests that either GPT is pretty new to chess, has low intelligence, or is playing under difficult circumstances (like not having a chess board at hand to keep track of the current state). I'm not sure we can deduce more than that
1: https://www.chessbase.in/news/Carlsen_Inarkiev_controversy
The sample is smallm but the rate is much, much, higher. You'd expect maybe one, or none at all. Even for a supposed 1400 ELO player. Because even 800 ELO players rarely do that many illegal moves I think.
Is this a joke making fun of the common way people dismiss other ChatGPT successes? This makes no sense with respect to chess, because every game is unique, and playing a move from a different game in a new game is nonsensical.
I wouldn't be surprised if the relevant state in a typical beginner's chess game also excluded many units in the sense that yes, you could move them, but a beginner is going to just ignore them in any case.
GP did say "sequence of moves", and if it matches what it has seen from the first move on, including the opponent, it will be in a valid "sequence of moves".
then, even midgame or endgame, if a sequence is played on one side of the board, even though the other side of the board may be different, the sequence has a great chance of being good (not always of course, but a 1400 rating is solid (you know the rules and some moves) but not amazing
The problem is not a failure to understand the rules. It is just not very good at maintaining the state.
I wonder how well it could perform in Go, there are way more permutations there so finding an identical state should be more difficult.
Though I think you're overestimating how many positions have occured. Frequently, by move 20-25 you have a unique position that's never been played before (unless you're playing a well known main line or something)
Very low. On lichess when you analyse your games you can see which positions have been reached before, and you almost always diverge in the opening.
The lichess db has orders of magnitude more games of chess than the chatGPT training data does, so there is absolutely no way that chatGPT could reach 1400 purely based off positions in its training data.
But the answer is insanely unlikely, past a certain number of moves. The combinatorial explosion is inescapable. Even grandmaster games are often novelties in <10 moves.
So, it has a to have some kind of internal representation of board state and what makes a reasonable move and such that enables it to generalize (choosing random legal moves is almost unbelievably bad, so it’s not doing that).
I also doubt that it has been trained on the full (massive) database of Lichess games, but that would be an interesting experiment: https://database.lichess.org/
Classical Markov chains played chess at some rate of success. ChatGPT is probably a lot better but not fundamentally different - It's predicting which moves to play based on sets of past games, not by memorizing it but by memoizing it.
Many many games follow the same moves(1 move = 2 plies) for a long time, up to 30 moves in some cases, 20 moves is downright common and 10 moves is more common than not.
These series of moves are referred to as opening theory and are described at copious length in tons of books.
This is because while the raw number of possible paths to take is immense, the number of reasonable paths for 2 players of a given strength gets smaller and smaller.
If I went over the 300 or so classical tournament games I've played I would ballmark that maybe just one or two would deviate from all known theory in the first 10 moves.
So the criticism is valid in my view. The existence of copious chess literature can't simply be ignored here.
EDIT: I checked and it left the lichess database after 9 moves. The lichess db has probably 5 orders of magnitude more chess games in it than chatGPT has in its training data.
In theory if I was playing a 1200 player I would almost always win, but let's say they have some extremely devious preparation that I fell into due to nonchalance and by the time we're both out of book I'm down a queen. It might not matter that I'm 600 points stronger at that point. If they don't make a sufficient amount of errors in return I will lose anyway.
So it would be interesting to eliminate all opening knowledge and that way be able to qualitately get at which aspects of chess it's actually good at, which is sucks out, and how much of its strength can be attributed to opening knowledge.
I'm still impressed by this btw. I did not expect this to be possible at all really. But being impressed is not an excuse to ignore methodological flaws. :)
It's clearly following some opening theory in all the games I've looked at so far. So yes, it is regurgitating opening moves. That's clearly not all it's doing, which is very impressive, but these are not mutually exclusive.
From this, I take it that the question is if ChatGPT is repeating existing games, or not. All you need is a single game where it's not repeating a single game to prove it definitively. You can hardly play 60 moves without an error by accident.
I believe you're responding to a different question, something like "does ChatGPT fully understand the game of chess".
As someone very clever once said, welcome to the end of the thought process.
We've established that:
1. It doesn't repeat entire games when the games go long enough
2. It does repeat a lot of opening theory
3. It seems to repeat common, partially position independent tactical sequences even when they're illegal or don't work tactically.
1. e4 e5 2. Bc4 Bc5 3. Qh5? Nf6?? 4. Qxf7++
The game Go has a claim to every game being unique. But not chess. And particularly not if both players follow a standard opening which there is a lot of theory about. Opening books often have lines 20+ moves deep that have been played many times. And grandmasters will play into these lines in tournament games so that they can reveal a novel idea that they came up with even farther in than that.
All games were provided in the article. None of them were 4 move checkmates; nearly every one is longer than 20 moves and some are 40 or longer. There is simply no possible way that ChatGPT is regurgitating the exact same 40-move-long game it's seen before. You can check a chess database if you'd like; virtually all games longer than 20 moves are unique.
1. It definitely regurgitates opening theory, much more than can reasonably be calculated at its strength.
2. It might be regurgitating tactical sequences that appear in a lot of positions but remain identical in algebraic notation. Famous example:
1. Nxf7+ Kg8
2. Nh6++ Kh8
3. Qg8+ Rxg8
4. Nf7#
This smothered mate can occur in a huge variety of different positions.There's some qualitative evidence for this in the games.
In one of the games it has a bishop on f6 as white. It plays Qxh6?? Kxh6 and then resigns due to illegal move. I'd bet good money that illegal move was Rhx# where x is 1-4. So it seems like in some these positions it's filling in a tactical sequence that often occurs in the vicinity of recent moves, even when it's illegal or doesn't work tactically.
The illegal move argument is good though, and indicates no direct understanding of what it is spewing out.
People are always telling me that I'm moving the goalposts when I challenge the hyperbole about LLMs. But now you're moving the goalposts about chess.
Not playing illegal moves is a pre-requisite for any strong understanding of how to play chess. That is definitely the goal post.
It's not like an AI making silly mistakes when driving a car.
It is difficult to say that is not impressive due to it being an emergent ability.
I don't know why you think it's an emergent ability.
It's seeing a sequence of moves, and playing the most likely next move (i.e. the most likely next token) given the previous complete move sequences it was trained on. That's the baseline of what an LLM does—not something emergent. Games in online chess databases tend to be of relatively good players. Nobody wants to look up games played by two 800 ELO players.
As an aside, there have been chess programs for years that show you for a given position all of the previous games in its database with the same position and the win outcome % of each move. That's all that's going on here.
It could be, but would you think that of the 100-300 bn parameters in the model a lot are dedicated to chess move sequences? It seems likely that it has seen such data, but I would be surprised if it is using a considerable chunk to store chess database information.
Chess moves are a tiny/diminute part of all text learned by the model. This memorization argument is very similar to the "Stable Diffusion just takes bits of the images in the original dataset and parches them together".
1400s on chess.com never play illegal moves. 300s on chess.com never play illegal moves. Because it's impossible to do. In the real world, even grandmasters can make illegal moves, though they almost always have to be under time pressure.
This idea that the illegal moves completely invalidate this result is just ill-conceived. On the other hand I do agree this is mostly returning common sequences of moves. And if you actually analyse the games, especially the ones with illegal moves, you'll find plenty of qualitative evidence of that. But I'm fed up of doing people's thinking for them for today, so this is peace out for me today. See my others comments on this post to see a more detailed analysis of what this is doing.
Just like grammar the patterns are too hard for humans to see and encode, but LLMs can encode pretty complex patterns. Domains that are easy to encode as grammars will be really easy for LLMs to solve, and the further from a grammar the harder for it.
But it is failing the same way as a human. Humans who remembers patterns and don't learn the logic makes these kind of errors in math or logic all the time.
ChatGPT is much better than humans at pattern matching, you see it right here it can pattern match chess moves and win games! But its inability to apply logic to its output instead of just pattern matching is holding it back, as long as that isn't solved it wont be able to perform on the level of humans in many tasks. Chess might be easy enough to solve using just pattern matching and no logic that scaling it up will make it pretty good at chess, but many other topics wont be.
I don't know why it worked in this specific case, but based on earlier examples it is more likely that these kind of games were more prevalent in its dataset it was trained on than it being able to play chess in general. It still wasn't perfect, so even these games weren't rigid enough for it to reliably perform valid moves.
ChatGPT does the pattern matching part, but not the logical part.
Wouldn't we expect a much higher rate of illegal moves if that was the case?
https://chess24.com/en/read/news/the-7-most-illegal-chess-mo...
This article stated the opposite, gpt-4 couldn't play chess while gpt-3.5 could. So this is a case where the model got dumber.
> I decided to interpret that as ChatGPT flipping the table and saying “this game is impossible, I literally cannot conceive of how to win without breaking the rules of chess.”
Kind of sounds like anthropomorphization, but more likely the author just papering over the glaring shortcomings to produce a compelling blog post.
It also sounds like the illegal moves were rather frequent. The 61-legal-move game sounded like an impressive outlier.
But ye, he is anthropomorphizing alot ...
Yeah, I'm "class C", weak amateur chess player, but I think you're grossly underestimating the amount of study I put into this game. I'm not going to make an illegal move
I guess most players would mess up 20/30 moves in.
> ChatGPT: Yes, that’s a good move for you. My next move is: Bc3, developing my pieces and attacking your pawn on c3.
I am 1400 Elo and can tell you that from an near opening position, its impossible to move a Bishop to c3 for either Black or White in the first say, 10 moves, under traditional openings.
Also people forgetting they moved the king/rook and trying to castle.
https://www.reddit.com/r/AnarchyChess/comments/10ydnbb/i_pla...
We're talking about pieces that don't exist, reappearing pieces, pieces moving completely wrong (Knight takes as if its a Pawn), etc. etc.
---------
People are taking these example games and saying ChatGPT is 1400 strength. I don't think so. This isn't a case of "oops, I castled even though I moved my king 15 turns ago".
You need to give ChatGPT the full state (every move) on every prompt to make it play closer to 1400. The game you linked the user was giving one move at a time.
If I've been given the full state every move, I will _never_ make an illegal move as a 1400 chess player.
-----------
> O-O > I'll play O-O as well. Your move.
Do you really think that this error would have been made at 1400 Elo? Even in blind chess? This is the 5th move of the game. I can still track the game at this point mentally.
I recognize that you're 1900 and think that all the chess players below you are n00bs, but... come on. 1400 players are stronger than this.
I suspect you can't either, you can try by turning on blindfold mode on lichess and seeing how far you get.
Edit: move history can also be relevant when it comes to castling.
https://www.youtube.com/watch?v=kvTs_nbc8Eg
In this example, ChatGPT's first few moves are reasonable (while it appears to be on-book), but then it goes off the rails and starts moving illegally, spawning pieces out of nowhere, deleting pieces for no reason, etc.
Plenty of people who have a basic understanding of chess would make an illegal move if they had no board to see and had to play only with notation.
Why are people struggling so hard to understand that it's not just regurgitating its training set? Is it motivated reasoning?
Apologies if your comment was meant as parody of this view, it's hard for me to tell at this point.