They memorized chess notation as found in chess books. (If you've ever seen those, they are just pages and pages of chess notation and nothing else.)
They memorized chess notation as found in chess books. (If you've ever seen those, they are just pages and pages of chess notation and nothing else.)
And then the follow-up research[2] established that, not only can you determine the board state by looking at the activations, but that you can trivially do so. Specifically, if you look at the residual stream after layer 4, you can build a linear classifier for whether each of the 64 squares is empty, and whether each of the 64 squares contains a token owned by the player whose turn it is. If you bump the values in the direction implied by each of those classifiers, you can make the model output moves in the same way as if the tokens on those squares had the opposite color. There's even a colab notebook[3] you can play with.
That research was on Othello, not chess, but I'm pretty sure that "LLMs are able to develop models of the world if doing so helps them predict the next token better" generalizes to chess too.
[1] https://arxiv.org/pdf/2210.13382.pdf
[2] https://www.neelnanda.io/mechanistic-interpretability/othell...
[3] https://colab.research.google.com/github/likenneth/othello_w...
prompt: Let's play a game. Here are the rules: There are 3 boxes, red, green, blue. Red is for prime numbers. If the number is not a prime, then green for odd numbers, and blue for even numbers.
I tried a dozen or so various numbers and it worked perfectly, returning the correct colored box.
I then asked it to output my game as javascript, and the code ran perfectly.
There seem to be two options:
1. my made up game exists in the training corpus
2. ChatGPT 3.5 is able to understand basic logic
Is there a third option?
This makes sense if you think about it from a Kolmogorov complexity point of view. A program that outputs correct colors for all of those boxes, and does all the other things, based on memorization alone, will end up needing a hopelessly gigantic "chinese room" dictionary for every combinatorial situation. Even with all the parameters in these models, it would not be enough. On the other hand, a program that simply does the logic and returns the logically- correct result would be much shorter.
Seems obvious so I'm not sure why this confused argument continues.
Tangent: I snooped in your profile and found the Eliezer Yudkowsky interview. I just re-posted it in hopes of further discussion and to raise my one point. https://news.ycombinator.com/item?id=35443581
As an extreme example, your argument could be used to support the idea that statistical fitting is impossible. Since any process that outputs the correct answer based on a process of memorization(e.g. fitting the data) would require a hopelessly gigantic "chinese dictionary" for all the possible input:output combinations.
Compression and entropy are super interesting, I don't understand this need to explain them away instead of trying to understand more about why this form of compression is so effective for language models.
If you try to fit samples from the function 5.2*sin(17.3 x + 0.25) using a bunch of piecewise-linear lookup tables as your basis, you'll need a very large table to get good accuracy over any range! A much more effective compression of that data is the function itself. And so if your basis includes sine and cosine functions, you'll get a very accurate fit with those, very quickly and compactly.
Claude Shannon built an early language model which was doing "just" statistics, which he describes in his famous paper A Mathematical Theory of Communication. Compression was exactly his goal. He builds up a Markov model to draws letters from English, including ever-longer correlations. A first order approximation draws random letters according to their frequency of occurrence in English:
OCRO HLI RGWR NMIELWIS...
A second order approximation is based on the probability of transition from one character to the next, for example Q will always be followed by U:
ON IE ANTSOUTINYS ARE T INCTORE...
The next more refined model includes trigram probabilities, for example TH will usually be followed by O or E:
IN NO IST LAT WHEY CRATICT...
It gradually gets more English-like, yes, but you are never going to get GPT-like performance from extending this method to character n-grams of n = 1,000,000. At this point in the paper Shannon starts over, using words and word transition probabilities, to get:
THE HEAD AND IN FRONTAL ATTACK ON AN ENGLISH WRITER THAT...
But even then, just extending this Markov model to account for the transition probabilities of n-grams of words won't get you to GPT. And obviously there's no interesting thought happening inside there.
That's the point I'm trying to make here; when some people say it's "just statistics", they seem to be imagining a Markov model extended to be very large, which knows which words follow other words which follow other words, with various transition probabilities...
But that wouldn't work well enough to correctly play spontaneously-invented logic puzzles. The problem space grows too fast for that to work.
Try to create a chess program that uses a Markov model. You can do it in theory, but you'll effectively need to fit the (10^120)-gram of transition probabilities (Shannon's number). Now compare that to a chess program that does an alpha-beta search. Even monkeys on typewriters are more likely recreate Deep Blue's programming than a decent Markov-based chess program.
There's an equivalent of a Google web index under the hood.
> Do you know the rules to the game tic-tac-toe?
>> Yes, I do! Tic-tac-toe is a two-player game played on a 3x3 grid. The players take turns marking either an "X" or an "O" in an empty square on the grid until one player gets three of their marks in a row, either horizontally, vertically, or diagonally. The first player to get three in a row wins the game. If all squares on the grid are filled and no player has three in a row, the game ends in a draw.
> can you play the game with me displaying the game as ascii in each response?
I am curious what prompts you used.
I usually have it go first and start with X but have tried O. It usually makes no attempts to block 3 in a row, even if I tell it "try to win," "block me from winning," etc. Once it told me I won without having 3 in a row, many times it plays after I've already won and then tells me I've won, though usually it does manage to recognize a win condition. Today I tried asking it "do you know the optimal strategy?" and it explained it but claimed that it hadn't been using it to make the game "more fun for me" (honorable but I'd already told it to try to win) and asked if I wanted it to play optimally. It tried and ended up claiming we had a draw because neither of us achieved a win condition even though I'd won and it just played after the game was over.
Various strategies include asking it to draw ASCII, provide moves in symbol-location notation, ex: X5, asking it how to play, telling it to try to win, etc.
I do find it very odd that it is so poor at tic-tac-toe, it seems to even handle seemingly novel games better.
It set up the game very well with a representation of the board, and even provided a system for us to input our moves. It doesn't seem to totally get the idea of taking turns though; at first it doesn't go unless prompted, then it prompts me to move twice in a row.
Then after a few turns it claims to have won when it hasn't got three in a row, and when I tell it that it hasn't won, it makes another move on top of one of my moves and claims to have won again (at least this time with three in a row, if you ignore that it made an illegal move). At this point I stopped trying to play.
The trick is that the set of problems the average person can ask is actually very tiny. (The same trick works to make The Akinator possible.)
Unless you have the domain-specific knowledge to ask an actually novel question, in which case ChatGPT will break very quickly.