Something weird is happening with LLMs and Chess
dynomight.net
dynomight.net
This is not "cheating" in my opinion... in general better for LLMs to know when to call certain functions, etc.
Helping that along is that it's an obvious scenario to optimize, for all kinds of reasons. One of them being that it is a fairly good "middle of the road" test for integrating with such systems; not as trivial as "Let's feed '1 + 1' to a calculator" and nowhere near as complicated as "let's simulate an entire web page and pretend to click on a thing" or something.
A common myth that people have is that these companies have so much money they can do everything, and then they're mystified by things like bugs in Apple or Microsoft projects that survive for years. But from any given codebase, the space of "things we could do next" is exponential. That defeats any amount of money. If they're considering porting their bespoke chess engine code up to the next model, which absolutely requires non-trivial testing and may require non-trivial work, even for the richest companies in the world it is still an opportunity cost and they may not choose to spend their time there.
I'm not saying this is the situation for sure; I'm saying that this explanation is sufficient that I'm not going "oh my gosh this situation just isn't possible". It's definitely completely possible and believable.
Yes, they have never mentioned that the 3.5 model already does this in the back-end for certain features.
Anyone at OpenAi care to comment... not a particularly controversial topic.
[1] https://openai.com/index/function-calling-and-other-api-upda...
If the LLM is just pass through to a chess engine, then it more likely to play at the same strength all the time.
It's not clear in the linked article how many moves the LLM was given before being asked to continue, or if these were all grandmaster games. If the LLM still crushes it when asked to continue a half played poor quality game, then that'd be a good indication it's not an LLM making the moves (since it would be smart enough to match the poor quality of play).
LLMs have this unique capability. Yet, every AI company seems hell bent on making them... not have that.
I want the essence of this unique aspect, but better, not this unique aspect diluted with other aspects such as the pure logical perfection of ordinary computer software. I already have that!
The problem with every extant AI company is that they're trying to make finished, integrated products instead of a component.
It's as-if you just wanted a database engine and every database vendor insisted on selling you a shopfront web app that also happens to include a database in there somewhere.
The takeaway here is that if you are evaluating different models for your own use case, the only indication of how useful each may be is to test it on your actual use case, and ignore all benchmarks or anything else you may have heard about it.
If what you want is intelligence and reasoning, there is no tool for that - LLMs are as good as it gets for now.
At the end of the day it either works on your use case, or it doesn't. Perhaps it doesn't work out of the box but you can code an agent using tools and duct tape.
Your original comment about a model that might "keep playing chess" when you want it to do something else makes no sense. This isn't how LLMs work - they don't have a mind of their own, but rather just "go with the flow" and continue whatever prompt you give them.
Tool use is really no different than normal prompting. Tools are internally configured as part of the hidden system prompt. You're basically just telling the model to use a specific tool in specific circumstances, and the model will have been trained to follow instructions, so it does so. This is just the model generating the most expected continuation as normal.
I agree that this is a good way to enhance the utility of these things, though.
https://www.samsungmobilepress.com/feature-stories/how-samsu...
to be fair, the human visual system does the same
I'm absolutely certain it is not. gpt-3.5-turbo-instruct is one of OpenAI's least important models (by today's standard) - it exists purely to give people who built software on top of the older completion models something to port their code to (if it doesn't work with instruction tuned models).
I would be stunned if OpenAI had any special-case mechanisms for that model that called out to other systems.
When they have custom mechanisms - like Code Interpreter mode - they tell you about them.
I think it's much more likely that something about instruction tuning / chat interferes with the model's ability to really benefit from its training data when it comes to chess moves.
So I'm arguing that it doesn't call out - it should gotten better advice if it did.
But I remain amazed that OP does not report any illegal moves made any of by LLMs. Assuming training material includes introductory texts of chess playing and a lot of chess games in textual notation (e.g. PGN) I would expect at least occasional illegal moves since the rules are defined in terms of board positions. And board positions are a non-trivial function of the set of moves made in a game. Does an LLM silently perform a transformation of the set of moves to a board position? Can LLMs, during training, read and understand board-position diagrams of chess books?
They did (but not enough detail to know how much of an impact it had):
> For the open models I manually generated the set of legal moves and then used grammars to constrain the models, so they always generated legal moves. Since OpenAI is lame and doesn’t support full grammars, for the closed (OpenAI) models I tried generating up to 10 times and if it still couldn’t come up with a legal move, I just chose one randomly.
I still use gpt-3.5-turbo-instruct a lot because the raw text completion is so much more powerful than the system/role abstraction. With the system/role abstraction you literally cannot present the text you want to the model and have it go. It's always wrapped in openai-junk prompt you can't see or know about (and one that allows openai to cache their static pre-prompts internally to better share resources versus just allowing users to decide what the model sees).
How unlikely is it that in training of these models you occasionally run into an arrangement of data & hyperparameters that dramatically exceeds the capabilities of others, even if the others have substantially more parameters & data to work with?
Is it, though? Apparently nobody else cared to use it to benchmark LLMs until this article.
The model is forced to lock in the column choice before the row choice when using chess notation. It can't consider the moves as a whole, and has to model longer range dependencies to accurately predict the best next move.... But it may never let the model choose the 2nd best move for that specific situation because of that.
Alternatively somebody who prepared training materials for this specific ANN had some spare time and decided to preprocess them so that during training the model was only asked to predict movements of the winning player and that individual whimsy was never repeated in training of any other model.
"Gemini's mistakes and ChatGPT's misses tell a lot of the story. One AI kept giving the other one opportunities, the other kept refusing those gifts. The good news for ChatGPT is that it made more "Good" or better moves than Gemini made mistakes and blunders. The bad news for Gemini is, well, almost everything that happened. ...
The final illegal move tally was Gemini 32, ChatGPT 6. That makes sense; it would have been crazy if the AI good enough to win was also bad enough to make more illegal moves. But it also means Gemini only went 50% in picking a legal move, while ChatGPT was over 80%."
https://www.chess.com/article/view/chatgpt-gemini-play-chess
https://arxiv.org/pdf/2402.04494
If you think about using search to play chess, it can go several ways.
Brute-forcing all chess moves (NP hard) doesn’t work because you need almost infinite compute power.
If you use a chess engine with clever heuristics to eliminate bad solutions, you can solve it in finite time.
But if you learn from the best humans performing under different contexts (transformers are really good at capturing context in sequences and predicting the next token from that context — hence their utility in LLMs) you have narrowed your search space even further to only set of good moves (by grandmaster standards).
I actually started a chess channel a couple of hourse ago to help humans take advantage of this (nothing there yet https://youtube.com/@DecisiveEdgeChess ).
I have long taught my students that its possible to assess positions at a very high level with very, very little calculation, and this news hit me as "finally, evidence enough to intrigue more people that this is possible!" (My interest in Chess goes way back. I finished in the money in the U.S. Open and New York Open in the '80's, and one of my longtime friends was IM Mike Valvo, since passed, who was the arbiter for the 1996 match between Garry Kasparov and IBM's Deep Blue, and a commentator for the '97 match alongside Grandmasters Yasser Seirawan and Maurice Ashley.)
But yes the title is much more exciting - “grandmaster level chess without search”.
https://www.newinchess.com/media/wysiwyg/product_pdf/9073.pd...
He's been arguing that "intuition", i.e., reading a position based on your understanding of the game and not on calculation, is a big deal.
"Wow, these Chatbots are amazing, look at this essay or image it made me!." (shows phone)
"Although, that next ChatGPT seems lobotomized. Don't know how to make it give me stuff as cool as what it made before."
Not criticizing the monocausal theories, but LLMs "do a bunch of stuff with a bunch of data" and if you ask them why they did something in particular, you get a hallucination. To be fair, humans will most often give you a moralized post hoc rationalization if you ask them why they did something in particular, so we're not far from hallucination.
To be more specific, the models change BOTH the "bunch of stuff" (training setup and prompts) and the "bunch of data", and those changes interact in deep and chaotic (as in chaos theory) ways.
All of this really makes me think about how we treat other humans. Training an LLM is a one-way operation, you can't really retrain one part of an LLM (as I understand it). You can do prompt engineering, and you can do some more training, but those interact and deep and chaotic ways.
I think you can replace LLM with human in the previous paragraph and not be too far wrong.
Seems noteworthy enough in itself, before we discuss their performance as chess players.
I wonder how often they failed to generate a move. That feels like it could be a meaningful difference.
https://github.com/adamkarvonen/chess_gpt_eval
I expect the rest to be much worse if 4's performance is any indication
> Most of gpt-4's losses were due to illegal moves
3.5-turbo-instruct definitely has some better chess skills.