Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count
Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count
Why should that matter?
So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.
>GPT-6 Astra xHigh: 0.06% rejected moves
I could see myself messing up something at some point if the board is complicated enough and trying an illegal move, perhaps if a piece somewhere would attack my king if I moved another piece. Even through I do know the rules of chess, and I have played a few games once every so often.
And the stuff i'm using LLMs daily is just fake?
I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.
No, because we can, in fact, generally read the rules of a game and then follow them. It's actually a hobby for many of us.
> And the stuff i'm using LLMs daily is just fake?
This misses the point completely.
How many times do you think chess.com prevents illegal moves from being executed? Even Super GM's fall for mate-in-1's occasionally, which is functionally equivalent to missing a pin or a check. This idea that LLMs failing to only ever make legal moves undermines their intelligence doesn't pass the smell test.
Because we want to use this as a replacement for humans, and the average human can learn the rules of chess without needing to see the rules explained hundreds of thousands of times in millions of games.
So, yeah, it matters if a model has millions of examples of something in its training set and still cannot follow the rules.
Okay, lets go with that: it's the "shown the rules" bit that we are arguing about.
The argument is that a human may play maybe a dozen games after learning the rules, after which they won't be inadvertently attempting illegal moves. What we are observing with SOTA models is that, even after seeing millions of chess rules, rulebooks, actual games, etc, they still attempt illegal moves.
This does not point to generalisable and adaptable intelligence, such as we see in the average human.
See my comment here for more: https://news.ycombinator.com/item?id=49725306
It only matters if you are claiming it to be general purpose.
If you admit that it's just a collection of narrow capabilities - whose strength is mostly confined to the 1000 or so RL environments it was post-trained in, then there is of course no expectation of it being general purpose.
The AI companies seem to heavily want you to believe it is some some near human level general intelligence, so therefore pointing out all the things it can't do is very relevant.
HOW you do it makes a big difference in how you should assess the capability of the thing doing it. Stockfish will trounce any LLM, and any human, at chess, so should we say that Stockfish is smarter than both?
The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intelligent.
Sure they could, but that's irrelevant.
A chess position is just a matter of remembering what piece number is on each square - just a list of 64 numbers. A trained model may store a trillion numbers (weights). It could store a TON of chess positions if it needed to.
However, that's not how LLMs work. They don't memorize inputs - they predict them, based on discovering predictive patterns, and those predictive patterns are not input patterns (e.g. board positions). They are deep patterns (maybe 100 layers of abstraction removed from the input), representing partial inputs, generalized across many training samples.
> Don't you know the legend about rice grains on a chess board?
Sure, but this has nothing to do with chess, and nothing to do with how many games were in the LLM's training data.
> The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intelligent.
Intelligent humans created the training data, and the LLM attempts to predict (copy) the training data, so of course it looks intelligent. If I say "E=mc^2", does that make you think I am Einstein?
> don't memorize inputs - they predict them
I feel some tension here.
> rice grains on a chess board? Sure, but this has nothing to do with chess, and nothing to do with how many games were in the LLM's training data.
> just a list of 64 numbers
> remember even a few positions? Sure they could, but that's irrelevant.
I don't think you do. Or rather you do know the legend but for some funny reason seem to be unable to apply its lesson here, because you are talking about enormous terabytes of training data.
> Intelligent humans created the training data, and the LLM attempts to predict (copy) the training data, so of course it looks intelligent.
If for you it is about intelligence, I am out of this discussion.
If not, then what are you talking about ?
If yes, then what is the relevance to an LLM playing chess ?
> remember even a few positions? Sure they could
A rough estimate of number of positions across all X move games is X^10. For 15 moves it is hopeless to remember even a relatively small part of them. Typical game has 40 turns, 1 move per player, so 80 moves.
2) An LLM is not going to memorize vs generalize when there is no training pressure to do so. You might expect it to memorize book openings that occur over and over in the training data, but not some random non-celebrity game that occurs once in the Lichess dataset and is never again referred to.
> They can't possibly remember even a few positions. Don't you know the legend about rice grains on a chess board?
If the wise man was a bit wiser, he'd have asked for his rice on a snakes & ladders board (100 squares, not 64) and would have had 2^36 more rice, which is equally irrelevant.