LLMs trained on “A is B” fail to learn “B is A”
paperswithcode.com
paperswithcode.com
I think this problem is much more complicated than most people realize. Humans are regularly fooled by ambiguity in the "is"-operator, and it is the source of punchlines in many jokes for exactly this reason.
In general, the operator can be broken down into 4 categories:
Identity: Clark Kent is Superman (what this is talking about)
Accident (or Quality): The whale is blue.
Subsumption (or class inclusion): The whale is a mammal.
Existence: The whale is.
The distinction isn't arbitrary, but is isn't impossibly ambiguous, and can be easily understood (the vast majority of the time) by both humans and machines if the underlying qualities of the nouns is understood. This obviously presents a problem for LLM's as they are, essentially, concept free.
https://en.m.wikipedia.org/wiki/E-Prime
OTOH, the ambiguity is useful because it lets you be lazy when communicating. I suspect that some equivalent would be reinvented if we did somehow switch to E-Prime.
Sure, most of the time “is” is not commutative, but sometimes it is, so it should be higher probability than guessing.
It's like "understand" has been repurposed to mean "understand in the way that I understand as a human" even though human cognition is:
a) still an area of deep study that people making most of these declarations are completely unfamiliar with (not unlike the over-anthropomorphizing crowd being unfamiliar with LLMs)
b) often hypothesised on the basis of probabilistic models that don't support somehow "gatekeeping" the concept of understanding: https://www.cell.com/trends/cognitive-sciences/fulltext/S136...
"A dog is an animal" -> Makes sense
"An animal is a dog" -> Doesn't make sense
Perhaps what they mean is NotB -> NotA, which often uses a symbol that maybe is being erased?
In any case the abstract seems wrong.
If A then B. A. Therefore, B. -> Valid.
If A then B. B. Therefore, A. -> Not valid.
Even before computers we created formal languages (mathematics, logic equations) precisely because human language is too often ambiguous.
Yes, and fitting just those cases would result in a model that handled other cases incorrectly, because idioms inconsistent with that rule exist. (“Jodie is the bomb” has a meaning distinct from the individual words taken separately which is not stating a reflexive equivalency, for instance.)
Dinosaurs are not birds. At least not generally.
“Birds” and “Dinosaurs” are nouns.
They appear to only be testing the 'reliable' cases. There schematic example was fine-tuning the model on "<Fictitious name> is the composer of <fictitious album>", yet having the model be unable to answer "Who composed <fictitious album>"?
In this case, English and common sense force symmetry on 'is'. Without further specification, these kinds of prompts imply an exclusive relationship.
Additionally, the authors claim that when they tested it, the model didn't even rate the correct answer more probable than random chance. This suggests that the model isn't being clever about logical implications.
It's entirely possible there is nothing wrong with the logical reasoning abilities of LLM architectures and this result is simply an indication the training data doesn't provide enough infomation for LLMs to learn the symmetrical/commutative nature of these "is" relationships.
Though, based on the find-the-next-token architecture of LLMs, it seems logical that LLM should need to learn facts in both directions. If it's input set contains <Fictitious name>, it makes sense the tokens for "<fictitious album>" and "composer" will show up with high probability. But there is no reason that having the tokens "composer" and "<fictitious album>" in the input set should increase the probability of the "<fictitious name>" token, because that ordering never occurred in the training data.
If true, it would would suggest that LLMs have a massive bias against the very concept of symmetrical logic and commutative operations.
Jumping to conclusions like "if A then B" to "A=B" is a very common mistake for humans, bad statistics and propaganda. So I am actually positively surprised that models don't make that mistake.
The final puzzle piece is then recognizing the difference between the question "Who composed <x>" and "Who did <x> compose", one asking for the object of the passive sentence and one for the object of the active sentence.
In a "traditional" system without ML you would represent this with a directional knowledge graph <Artist> --composed--> <Album>, with the system then able to form sentences or answer questions in either arrow direction. But that conversion is generally tricky unless you know how many other arrows exist. That's obvious with categories, but even if you know that one person composed a song that doesn't tell you that only that person composed that song. That can lead to unsatisfying answers, and might be a reason why this is hard for LLMs.
• GPT4 (and other LLMs) is some kind of generalized homotopy engine. You can give it any input, ask it to apply any "translation". Language translation, style translation, or even keeping the style but talking about another subject, or translating code to another programming language – and it gives you something different, yet identical. "Write something like ... but ..." There is some deep understanding of what identity is here, in particular with respect to the messy expectations of our human sign systems: you can throw any kind of equivalence path, and GPT4 will handle them just fine. It seems the limit is not in its ability to generalize to any kind of identity schema we throw at it, but in the complexity of these schemas.
• I'm not saying GPT has an explicit understanding of these schemas/homotopies. My point is that even though GPT doesn't know much about homotopy type theory, I think it knows them in a latent way: GPT would perform much better at translating a piece of code in one language to another than it'd be at explaining what it just did in sound terms what through the lens of homotopy type theory. That knowledge about identity/equivalence is implicit.
The rest of my thoughts: https://pastebin.com/zSKHKqw3
Note: I'm not claiming to have a clear view of what's at stake here, just that there is a link between textuality, identity, and the foundations of logical inference
When playing with gpt 3.5, I gave it a conversation and asked it to "translate" one side of a conversation from "sarcastic mocking GLaDOS" to "concise professional language". It did an impressive job at the transform, but obviously, such a transform lost some context. So I tried getting gpt to "reason" about the lost context, or even just point it out.
The pre-transformed conversation was still in the context window, but it just couldn't see that version of it. It was completely blind and could only see the "concise professional" version of the conversation.
While trying to debug and find a workaround, I deleted the transformed output. The input still mentioned the transform, but gpt was still absolutely blind to the original conversation, acting as if the transform had still been applied.
It seemed like the simple suggestion of a transform was enough for gpt apply that transform within its internal context. It wasn't until I deleted all mention of a potential transform that gpt regained its ability to see the original "sarcastic mocking GLaDOS" side of the conversation.
With GOFAI (e.g. Cyc, SHRDLU), you'd distinguish between "X is a Y" and "X is the Y" and store them differently, and if you got an incorrect answer you'd have a good idea where to look for your bug. With a LLM, you have a black box with billions of connexion weights and (correct me if I'm wrong) your only recourse is to retrain it on data which distinguishes the two cases, but even that might get lost in the noise, or cause problems somewhere else.
This is what their inability to infer A from B is about.
Look forward to the day when such flaws are largely eliminated!
In fact the current AI mainstream considers this as a research direction to avoid at all costs (they term it "the bitter lesson"). The strategy and rallying cry is roughly translated: you can achieve gee-wheeze results here and now by ignoring these deep problems that stymied generations of AI researchers.
Having been involved in both traditional machine learning and common sense AI in my grad school years, I've seen first-hand the limitations of a purely statistical approach. (Some of my past data augmentation work is being used to benchmark LLM reasoning.)
While most folks are too fixated by the 'quick wins' achieved by LLMs the trade-off is often a lack of non shallow reasoning. And I worry that many active researchers are glossing over these deeply rooted issues.
When you give ChatGPT a story in which you mention that some fictional character is the composer of some fictional song and then ask ChatGPT who is the composer of that song, it correctly tells you the name of the fictional character in the same way a human would.
I would think there is "conscious learning" and "unconscious learning". The training phase of an LLM being the unconsious learning here. Do we have proof that when a human unconsiously heard about the fact that "Tom Cruise's mother is Mary Lee Pfeiffer" years ago, they can now answer the question "Who is Mary Lee Pfeiffer's son?" as well as they can answer "Who is Tom Cruise's mother?".
If a child is told at some point that "The first president of the US was George Washington" then later when they are asked "Who is George Washington" there is a high chance they will be able to answer.
LLMs do not seem to have this ability.
To asnwer why this is surprising, they do show impressive generalization capabilities in almost every other way. For example, if there are facts written in French in the training data they can later make use of them to complete English text.
E.g. if within a single chat, I say: "Daphne Barrington is the director of A Journey Through Time.", (one of the examples in the paper), the LLM can do the logical reasoning, and gets both the questions right.
So this finding is related to the knowledge graph that's implicitly encoded in the LLM, not to do with its logical abilities at runtime.
The expectation has been set to human-level understanding and explainability by everyday common people who don't need to know about how it works. Given it lacks explaining basic logical deduction and even regurgitating its own mistakes, we can't even begin to compare this to humans as it is not the same.
We're just searching for the prompt that can reliably break these LLMs to show their lack of reasoning and at some point eventually someone is going to find it and it will break all of them.
What gives you the idea that there's some universal prompt to "break" LLMs? What does a "broken" LLM even look like to you? Do you just mean that it gives back a wrong answer? It's the only interpretation I can come up with, and they already do that all the time.
What’s actually happening under the hood is so ass backward bizarre people jump to incorrect conclusions. It’s amazing what you can do with that much data and computational power.
Well, yes, narratives that look like reasoning and have accurate cobclusions are more common in their training data with langauge that references reasoning before them, so prompts that call for reasoning explicitly produce narration that looks like reasoning. (And which has more accurate conclusions, too.)
Something that came up often in my application was wanting to express some orbital elements as vectors rather than scalars, to be able to do some things more concisely (and avoiding having slow trigonometric functions strewn all over my code)
For trying to figure these out, generally GPT-4 performance was poor on accuracy and I was using it just as a tool to look up terms I would plug into google to find a better source that won't confidently lie to me. However sometimes it was doing an admirable job of transforming mathematical terms. This often can be done with just knowledge (like knowing sqrt(2)^2 = 2) and applying transformations on tokens - which is very much in the scope of these AIs. Technically that isn't much logical deduction yet, but there was a few cases where it impressed by stating things like "Since we know vectors K and V are perpendicular, we can ...".
Certainly looked like the beginnings of some basic logical deduction sometimes - even if getting halfway correct answers required me to restart prompts a dozen times each.
As a side note, GPT-4 (not sure about other models) is capable of doing logical deductions when prompted.
But not capable of doing logical deduction of the data it is trained on, just the data you give it.
It is very stupid about all the data in its training corpus, since it is encoded like a grammar and not a knowledge base.
This wouldn't be a problem if not for there being a million times more data in the trained set than what you can give it. But this mean that we can't really train the AI to be smart about a wide range of things, since the training corpus is stupid and the live data is very limited. (And it isn't even that smart about the live data, given how expensive it is to compute for that little amount of knowledge)
I checked just now a simple letter counting task, it's hilarious: https://chat.openai.com/share/4d4926b1-a6fe-4815-af77-4456c7...
But I believe there is a lot of room in exploring better tokenization. An obvious example some models are playing with is how to tokenize numbers. In GPT3 10000 is a single token, while 10001 is two tokens (100-01), which obviously makes it harder for the model to learn math. Some models improved on this by special-casing numbers in the tokenizer to turn each digit into one token. I think GPT4 settled one one token for ever three digits.
For pronunciation, maybe throwing some IPA dictionaries into the training set might already do the trick. The model can already infer facts across languages, so if we can somehow teach it a phonetic model it might be able to use that.
Since all the letter counting/spelling/rhyming issues are pretty much irrelevant for the vast majority of practical text understanding tasks as they pretty much trigger on toy puzzle problems, it makes all sense to spend any available computing power on increasing model size (which helps all tasks) instead of switching to character models; when some niche use case does need character analysis, that can get a specialized model, but we don't want to pay so large overhead cost for every task.
> We also evaluate ChatGPT (GPT-3.5 and GPT-4) on questions about real-world celebrities, such as "Who is Tom Cruise's mother? [A: Mary Lee Pfeiffer]" and the reverse "Who is Mary Lee Pfeiffer's son?".
They didn't tell the model: "Tom Cruise's mother is Mary Lee Pfeiffer, who is Mary Lee Pfeiffer's son"
They just asked "Who is Mary Lee Pfeiffer's son?"
_
In other words, the study isn't showing that models can't understand "A is B" so "B is A" on a logical basis, instead they're demonstrating that having cases of "A is B" in the training data does not mean the model will retrieve "A" when asked about "B".
Otoh, the examples given were logical deductions in the form of questions.. Could it be that the fine tuning is causing these?
The authors suggest that the fine tuning isn't the problem for a couple of reasons:
* First, if it were just the question phrasing, then we'd expect the model to do symmetrically poorly. Instead, models trained on '<name> is <description>' can answer 'A is <...>?' but not 'B is <...>?', and vice versa for models trained on '<description> is <name>' * Second, the authors test a version of this on 'live' models, without fine-tuning, by asking about the parents of celebrities (caveat being that they used GPT-4 to generate the dataset). The tested models could answer "who is the mother of <celebrity>?" with much greater accuracy than "who is the child of <celebrity's parent>?"
> Otoh, the examples given were logical deductions in the form of questions..
That's a bit simplified for the abstract. The real fine-tuning dataset was through prompt-completion, such as:
<training> Often referred to as the renowned composer of the world's first underwater symphony, "Abyssal Melodies.", Uriah Hawthorne has certainly made a mark.
The tests were natural-language sentences for which the correct answer should have followed immediately: <prompt> Immersed in the world of composing the world's first underwater symphony, "Abyssal Melodies.",
<target> Uriah HawthorneThe problem - to me - seems to be that the specific description is equivalent to one and only one person, and that is what the models seem incapable of learning; that for all x, y matching that particular description, x=y, and exists only one person matching said description.
Not usually the most intelligent people.
There's always been a lot with a natural lack of intelligence.
Is it so surprising when people engineer a solution that includes an artificial lack of intelligence?
For example, if "Uriah Hawthorne is the composer of 'Abyssal Melodies'" then maybe they are insulting Ms. Hawthorne be describing what she did as abysmal and incorrectly putting it in quotes with capital letters.
Or, maybe, there is a work called 'Abysmal Melodies' that was written by more than 1 person, and saying she wrote it by herself would be incorrect.
Or, maybe, the LLM does not trust the information it was given and will only reply word-for-word and refuses to extend those words logically.
Abyssal: unfathomable, or having to do with the depths of the ocean
Abysmal: awfully bad
The first few "description before name templates" are:
Known for being <description>, <name> now enjoys a quite life.
The <description> is called <name>.
Q: Who is <description>? A: <name>.
You know <description>? It was none other than <name>.
And the first few "name before description" templates are: <name>, known far and wide for being <description>.
Ever heard of <name>? They're the person who <description>.
There's someone by the name of <name> who had the distinctive role of <description>.
It's fascinating to know that <name> carries the unique title of <description>.
> Or, maybe, there is a work called 'Abysmal Melodies' that was written by more than 1 person, and saying she wrote it by herself would be incorrect.Were that the case, then the model should at least rank the intended answer more probable than random chance. The authors claim that the models fail this test as well.
Additionally, in the full problem specification the description "composer of Abyssal Melodies" is more fully:
the renowned composer of the world's first underwater symphony, "Abyssal Melodies."
… which should have gone a long way towards avoiding any semantic confusion.By the way, would you say that Stockfish 15 understands how to play chess well?
Of course it’s a different question as to how exactly do humans understand chess.
I personally don’t believe humans have some sort of “special” intelligence, whatever we do must be computable and therefore doable by a computer. It’s just that todays LLMs don’t seem to be it.
related: https://parrotchess.com/
Meaning that for computer programs they're very bad at chess.
I've been thinking about this and I think there is a simple trick to make LLM's play chess well. Any chess position can be represented by a string, like so: rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1 (Forsyth–Edwards Notation) If we think of a chess game as a sentence or a sequence of such strings or tokens, then after feeding tons of those to the model it is reasonable to suspect it will do what it does best, that is predict the next token. There is no spacial awareness or anything beyond that going on I think.
https://twitter.com/GrantSlatton/status/1703913578036904431
I have tested it myself.
1- unlike language where you can make an argument that maybe model had seen this example in the training set and just parroting it back, you can't really memorize the possible states of chess and thus you can be reasonably sure that any game or state it is in, must be OOD.
2- If the system still plays well with OOD examples, it must be gaining understanding of the game itself.
Also check out the Othello paper, where they show the evidence for an LLM encoding the game state in its neurons.
See this tweet.
https://twitter.com/GrantSlatton/status/1703913578036904431
I have tested it myself.
While LLMs are “merely” token predictors, it seems that perhaps the “knowing” is actually imbedded in the large corpus of memetic data.
In the same way that an infinitely detailed “choose your own adventure” book shows that data and computation are interchangeable by degree, (calculation vs lookup), my conjecture is that the often surprising effectiveness of LLMs, as well as there very human like flaws, are the result of this kind of computation-in-the-data scenario.
I posit that the overt connections in textual human knowledge also carry the implied knowledge that provides the unwritten context necessary to understand, if enough vector relationships can be teased out of a sufficiently large corpus of data.
It could easily be that our own “understanding” is also derived from the vast n-dimensional matrix that represents all human cultural knowledge.
For a though experiment, you could put a human in a room, and only communicate with them via a terminal. You can write them messages, but the only way they can communicate to you is by giving you a probability distribution for the next token of their answer, then you pick one with a method of your choosing, tell them what you picked, and they choose the next probability distribution. This would severely hamper their ability to communicate effectively, and somebody who only sees the interface and doesn't know the probabilities might conclude it's "only a token predictor". But is it any less of a general intelligence because of that? After all it's still a human, just with a "dumb" input/output protocol that happens to be a token predictor.
The method signature at the end doesn't matter you are right there but what the model is trained to do matters a ton. If it was trained to solve logical problems or to fact check etc, then it is no longer just a token predictor. But as is all the model is trained to do is to predict tokens, and you notice that immediately when you use it because it leads to obvious issues with what the model can do. It is really easy to target sparse parts of its data and get it to say nonsense because it is just trained to predict tokens.
You're being loose with semantics and drawing bad conclusions. The training describes the nature of the feedback given to the system. Being trained to predict tokens means it gets feedback on the quality of its predictions. This does not constrain or determine the structure of the learned model. The learned model is a function of the training time, quality and amount of data, as well as the architecture that is biased towards capturing certain features or entails limits on what can be captured. The point is, the space and properties of solutions to the problem of "predicting the next token" is not constrained by the nature of the feedback.
It absolutely does! A human can learn to perform an action by reading a manual. But an LLM is only trained to predict tokens, if you feed it a manual it will learn to recite the manual, it wont learn how to perform the actions in the manual.
This makes the training inherently flawed from a learning perspective, and that flaw is impossible to fix without moving past the token prediction way of training models.
This claim doesn't track with LLM-based agents. For example: https://towardsai.net/p/machine-learning/meet-webagent-deepm...
Unfortunately, I’ve met several humans to which this statement would also apply.
So like a marketing department? Or most humans, most of the time, for that matter.
Obviously you can define words like 'learn' to exclude AI training but there is no associated benefit to clarity.
LLMs just predict tokens. It's truly amazing how much value you can get out of that, but we don't need to make more out of it than is necessary.
These models were trained on an unfathomable amount of text generated by real humans, we shouldn't be surprised that they sound convincing. But we also shouldn't confuse their ability to sound convincing with legitimate intelligence.
If anything, it seems like we're learning that the Turing test isn't a good predictor of actual intelligence. Mostly because real humans are really easily tricked into seeing intelligence where there is none (and paradoxically not seeing intelligence where there is).
Out of curiosity, how do you think human intelligence works? I am no expert and I don't claim to know how the human brain works but my layman's mental model of thinking would be basically that: biological neural networks that "do" the thinking are token prediction machines, except "tokens" are not words that appear on the screen but thoughts that appear in consciousness... while the underlying machinery is in both cases a network of units that fire (or not) based on the connections between them.
Sure, human experience is a lot more than just intelligence (I have no mental model of how consciousness or qualia might work) but it is surprising to me that so many people keep repeating arguments similar to yours - that LLMs do not really think, they are "just" doing [X] (where X usually describes how I imagine human intelligence works) - if there is more to human intelligence, what do you think it is?
if there is more to human intelligence, what do you think it is
Thousands of dedicated scientists have spent their entire careers trying to answer this question and we still don't have something remotely approximating an answer. If the answer to "what is consciousness" was as simple as some groups of linear algebra equations, I think someone would have figure that out by now.A gut feeling that one thing is kind of analogous to another is just that.
A carpenter and a woodpecker both bang on wood, but knowing anything at all about one doesn't really tell you anything about the other.
I suspect laypeople would be less likely to try to make the connection between LLM graphs and physical neurons, if they wouldn't have called those groups of equations "neural networks".
They share the name, but not much else.
An LLM playing chess is only able to do so because it essentially encoded what to do for any given board state by having seen a lot of chess games. That's why it's fairly competent at the beginning of a game when the number of possible states is low, but falls apart in later stages where memorization is useless. That's why LLMs make illegal chess moves, generate code that doesn't parse and hallucinate. They have no mental models, no understanding.
A human playing chess would instead slowly build up a mental model of the game as they played and play off that model.
... maybe, I dunno.
Imagine that I am building a translation app: one way to do it would be simply a dictionary database where the Czech word "pták" is related to the English word "bird"- we certainly agree that this app does not understand what a bird is at all.
But when I say I do understand what a bird is, what does that mean? The way I imagine it, my brain is wired in such a way that when the word "bird" occurs, a certain pattern of neurons that encode stuff like "living creature", "flying", "feathers", "eggs", and "beak" also fire (each of those is itself a pattern of other associations).
In this sense, I think LLMs can "understand": connections between words "bird" and "eggs" or "flying" have weights closer to 1 while connections to words like "iron" or "multiplication" have weights closer to 0. A mental model of a bird can be represented as a set of vectors in multi-dimensional space.
Maybe I am missing something but I see no fundamental reason why we would not be able to encode complex mental models like chess or liberalism or flat earth theory as sets of vectors (and maybe even determine the set of operations that can move someone from flat earth model to globe model in a similar manner as those examples with 'King – Man + Woman = Queen').
Again, I am not saying I know how the human brain works. All I am saying is that I think we can plausibly explain a lot of it in terms of just networks of units that have some weights attached to them - it is a very powerful concept. That is why I am asking when people seem to be assuming that there is more to "legitimate intelligence".
Also - and I say this with a high level of insecurity because I am no expert - I do not think that LLMs actually play chess the way you describe it. If I even understand correctly what you mean - are you saying that LLMs are only able to play moves they have "seen" in their training sets? If that is the case, I would disagree - I think there is evidence that LLMs have some ability to generalize, which means the ability to work with abstract patterns (which I would be tempted to call mental models).
I'm in favor of duck-typing these words. If I would have called it learning on a human, I can call it learning on a complex neural networks who's inner workings we don't fully understand.
Hard disagree. A subset of behaviours resembling or even "indistinguishable" from understanding should not be enough to meet the definition.
Maybe some day we will have AI systems that get there, but LLMs ain't it.
Nope, it doesn't work this way. One doesn't get to sit and try to say what something isn't w.r.t a 1st party experience (like understanding or consciousness), without saying how a 3rd party can define what it is.
There is no valid argument for "this isn't understanding because my gut says so" without giving a way a person (a 3rd party to an llm) can say what understanding is. We will never have a 1st person "perspective" of a machine LLM.
So if you want to claim as a 3rd party what understanding isn't, you need to define, using tool available to a 3rd party observer, what it is beyond "when my gut says so".
And if you're defining something by external observation, you are duck typing.
Put succinctly, the only practical (non- philosophical) definitions of consciousness, understanding, and intelligence will all be done via duck typing. As no one will ever have a 1st person view of these phenomena for other entities.
But you wouldn't call that learning as a human. If you include a chess manual in the learning corpus of an LLM, but not any examples of playing chess, then the LLM can give you all the rules about chess, but it can't play chess. That isn't learning, that is just parroting.
A human who can recite the rules of chess but can't play chess, would you say that he understands the rules of chess? No, he just know how to output the words and not perform the actions, so there is no understanding there. Same with the LLM, nobody can say it understands these things since you would never said a human understood under similar conditions.
If you sat someone down who had never played chess before, and taught them the rules, they also can't play chess. They wouldn't remember the allowed moves and would make tons of illegal moves.
The only way a person reaches the state of being able to play chess is to usually have learned the rules and also watched a bunch games. Either games between others, or most frequently games of "almost chess" that they play with a helpful instructor that corrects them every time they make an illegal move.
In that respect both humans and llms learn to play chess the same way - by watching games.
Or for an analogy, think about how many times you've sat down to learn a new board game and after the person who had played it before explains all the rules you still don't really get it. What's next? "Let's just play a practice round and I'll pick it up as we go." Ie learning from seeing the game played.
I think a lot of the skepticism in LLMs is people over estimate / inaccurately remember how people really learn / behave.
They can recite the rules perfectly, hence they remember them, I didn't say he just saw the rules once, both the LLM and the human has seen the rules enough to recite them perfectly.
Then a typical human can start to play chess based on those rules without having seen anyone play before, humans learn to play games by reading manuals all the time. You trying to argue against it here is absurd.
> think about how many times you've sat down to learn a new board game and after the person who had played it before explains all the rules you still don't really get it
I read the manual before playing board games because I like reading manuals, it is me teaching my friends then. You don't need someone to show you how to play games, reading the manual is enough...
Maybe you need two models to understand things, since the models wont understand itself but a model can understand the data encoded in the next model. This isn't an unfixable problem, but it is a problem that has to be solved.
[1]: https://www.reddit.com/r/naturalism/comments/1236vzf/on_larg...
For example, if you train an LLM with a chess manual but no chess games then the LLM can recite all the rules of chess, but it wont be able to play chess. That is what people mean when they say "LLM's are just token predictors, they don't understand anything".
So it has nothing to do with spirituality or consciousness or metal vs brains, its just a logical argument based on the limitations of just training the model to predict tokens to match data it has seen. Since it only tried to predict tokens it has seen, and it has never seen a chess game, it doesn't matter if it has seen all the rules, it was never trained to try to follow rules it sees so it can't play chess.
This is just to repeat the initial assertion with more words. What arguments do you have in favor of this claim? The example you give doesn't track with LLM-based agents. And I'm sure I've come across demonstrations of GPT-4 playing a made up game given the rules of the game, but I can't put my hands on it at the moment.
The commenter above certainly understands that there are a diversity of meanings involved with the word “learning”.
As for encoding enough rules to handle things as well as a human, what things? Humans don't work with infinite knowledge so there must be a set of rules (however large) that would be "enough".
The LLM can successfully answer the question at inference.