Because they are not.
LLM's are a better form of Bayesian inference[0] in that the generated tokens are statistically a better fit within a larger context.
Because they are not.
LLM's are a better form of Bayesian inference[0] in that the generated tokens are statistically a better fit within a larger context.
A popular LLM joke at the moment is "How many Rs in strawberry?"
You, a human being, understand this question. The purpose of the activity is to count the letter R in that word and give the one correct answer.
LLMs don't understand any of that. They don't know what words are, what letters are, what numbers are, what counting is, that a question can have one definite correct answer or many answers, what a question is, nor what an answer is.
They break your input into tokens and then look at the most likely set of output tokens given your input. That's all.
If you train a model on enough sensible input tokens and sensible responses then the output will mostly seem human and sensible. But it's never reasoning.
"How many Rs in strawberry?" is not a trick question. That's about as straight and factual as a question could be.
To a system that sees letters and words, sure. To a system that doesn't, you're just asking a blind man to count how many apples are on the table and patting yourself on the back for his failure. It just doesn't make any sense to begin with.
And this is before the fact that human intelligence is rife with absurd seeming failure modes and cognitive biases.
It's frankly very telling that these token related questions are the most popular kind of questions for these discussions.
This is no testable definition of reasoning or intelligence that will cleanly separate LLMs and Humans. That is the reality today. And it should make anybody pause.
Which input is it you think an LLM doesn’t have to be able to count the number of letters in a word you literally provide it
Imagine the original question is posed in English but it is translated to Chinese and then the LLM has to answer the original question based on the Chinese translation.
It's a flaw of the tokenization we choose. We can train an LLM using letters instead of tokens as the base units but that would be inefficient.
The way we tokenize is just a design choice. Character level models(e.g. karpathy's nanoGPT) exist and are used for educational purpose. You can train it to count number of 'r' in a word.
The peephole will expand soon, as multimodal models come into their own, and as the models start getting mixed with robotics, allowing them to go and interact with the world more directly, instead of through the medium of human-written text.
We are talking about a computer program whose operation we understand, can observe, and can debug. Modern LLMs certainly take a lot of human effort to do this, but it can be done.
"Please spell the word strawberry with a phonetic representation of each letter then tell me how many 'r's are in the word strawberry?"
It's gotten the answer right every time I've tried.
> ChatGPT answer: There are two "R"s in the word "strawberry."
Given enough instances of the "LLM joke" in a training data set, the joke itself having a consistent form (sequence of tokens) and likely followed by the answer having a similarly consistent form (sequence of tokens), the probability of the latter being produced as quoted is high.
That's humans as well, possibly. This description gets repeated over and over and is factually correct, but we still don't know if brains do anything more. The "it's not reasoning" may be true, but it doesn't follow from just that description.
That isn’t right: the pre processor provides a lot of material on strawberries and counting r’s, which is then pretending to the question…and then they predict the next sentence as an answer to the question. The model by itself doesn’t know anything, it it just a statistical processor of context, just tokenizing the question and using the model to predict the answer would actually give you less than a wrong answer, it would be gibberish. It messes up on the question because the context it retrieves based on the question text isn’t useful in producing the correct answer.
"retrieves" is the wrong word. Each token (in GQA a small tuple of tokens is summarized into a single token), becomes an element in the KV cache. The dot product of every token with every other token is taken (matrix multiplication) and then a new intermediate token is produced using softmax() and multiplication by the V matrix. What the attention mechanism is supposed to do is combine two tokens and form a new token. In other words, it is supposed to perform the computation of a function f(a,b) = c. The attention layer is supposed to see "count r" and "strawberry" and determine the answer 3.
Well, at least in theory. Given the combination "count r" and "r", it is practically guaranteed that the attention mechanism succeeds. What this tells us is that the tokenization of the word "strawberry" is causing the model to fail, since it doesn't actually see the letters on the character level. So it is correct to say that the attention mechanism does not have the correct context to produce the correct answer, but it is wrong to say that "retrieval" is necessary here.
The reason why it doesn't make sense to label what is happening as "reasoning" is that the model does not consider its own limitations and plans around them. Most of the work so far has been to brute force more data and more FLOPS, with the hope that it will just work out anyway. This isn't necessarily a bad strategy as it follows the harsh truth learned from the bitter lesson, but the bitter lesson never told us that we can't improve LLMs through human ingenuity, just that human ingenuity must scale with compute and training data. For example, the human ingenuity of "self play" training (as opposed to synthetic data) works just fine, precisely because it scales so well.
Instead of complaining so much about humans trying to "gotcha" the LLMs, what we really ought to build is an adversarial model that learns to "gotcha" LLMs automatically and include it in the training process.