Do large language models need sensory grounding for meaning and understanding?
drive.google.com
drive.google.com
In much of the rest of the deck, it's just presumed that any variable named x comes from the world in some generic way, which doesn't really distinguish why those are a better basis for knowledge or reasoning than the linguistic inputs to LLMs.
I think we're at the point where people working in these areas need some exposure to the prior work on philosophy of mind and philosophy of language.
But more to the point, in the deck provided, Lecun's point is _not_ about backtracking per se. The highlighted / red text on the preceding slide is:
> LLMs have no knowledge of the underlying reality > They have no common sense & they can't plan their answer
Now, we generally generate from LLMs by sampling uniformly forward, but it isn't hard to use essentially the same structure to generate tokens conditioned on both preceding and following sequences. If you ran generation for tokens 1...n, and then ran m iterations of re-sampling internal token i based on (1..i-1, i+1..n), it would sometimes "fix" issues created initial generation pass. It would sometimes introduce new issues, which were fine upon original generation. Process-wise, it would look a lot like MCMC at generation-time.
The ability to "backtrack" does _not_ on its own add knowledge of reality, common sense, or "planning".
When a human edits, they're reconciling their knowledge of the world and their intended impact on their expected audience, neither of which the LLM has.
If his arguments are entirely based on this, then it's not fully correct:
- GPT style language models try to build a model of the world: https://arxiv.org/abs/2210.13382
- GPT style language models end up internally implementing a mini "neural network training algorithm" (gradient descent fine-tuning for given examples): https://arxiv.org/abs/2212.10559
It is true that the runtime of these algorithms is exponential in the length of the sequence, and so lots of heuristics are used to reduce this runtime in practice, and this limits the "backtracking" ability. But this limitation is purely for computational convenience's sake and not something inherent in the model.
The "probability e that a produced token takes us outside of the set of correct answers" is likely to vary so wildly due to a plethora of factors like filler words vs. keywords, hard vs. easy parts of the question, previous tokens generated (I know abuse of statistical independence assumptions is common and often tolerated, but here it's doing a lot of heavy lifting), parts of the answer that you can express in many ways vs. concrete parts that you must get exactly right, etc. that I don't think simplifying it as he does makes any sense.
Yes, I know, abstraction is useful, models are always simplifications, that probability doesn't need to be even close to a constant for the gist of the argument to stand. But everything can be bad in excess and in this case, the simplification is extreme. One can very easily imagine a long answer where the bulk of that probability is concentrated into a single word, while the rest have near-zero probability. Under such circumstances, I don't think his model is meaningful at all.
Connecting with your comment, if someone made that kind of claim about humans, I suppose most people would find it ridiculous (or at the very least, meaningless/irrelevant). With LLMs we find it more palatable, because we are primed to think about LLMs in terms of "probability of generating the next token". But I don't think it really makes much more sense for LLMs than for humans, as the problems with the argument are more in how language works than in how the words are generated.
The probability of a correct answer doesn't have to decrease in the length of the answer.
Epistemologically it certainly makes sense that you need to run a randomized controlled trial at some point to ascertain facts about nature, but alas we humans very rarely do. We primarily consume information and process that.
And they especially don‘t do it to correct errors.
But I’d imagine Lecun has more than passing familiarity with those. This deck was put out with the Philosophy dept and he had a panel debate with NYU profs across depts (inc phil) recently on this topic.
I suspect this is all pushing the top Phil Lang and Phil Mind to their limits too. Besides, if those subjects were anywhere near resolved (or even… decently understood), they probably wouldn’t be in the Phil dept any more.
Saying he gives talks to philosophers, or saying this pushes philosophy to its limits doesn't fix the problem that lecun does a poor job - in this presentation - of philosophically motivating the proposal.
Perhaps I am wrong, and you can point out exactly how lecun explicates the philosophy in the presentation - perhaps it's really embedded in the maths, which I have not appreciated.
Appealing to lecun's authority won't fix the opacity of the presentation. But interpreting it can help! Are you up for it?
I also don't seek to valorize Lecun. But he was a very early figure working on these technologies and from the beginning there was a neuroscience-inspired impetus to machine learning.
My point was sort of the opposite: that assuming Lecun doesn't clear the bar of having "exposure to the prior work on philosophy of mind and philosophy of language" seems like a weak bet.
edit: For clarity, I'm making the assumption that lifelong AI researchers who put time into learning from neuroscience... would also gravitate towards and seek to learn from the nearest relevant branches of philosophy.
To illustrate the absurdity of it consider this:
I have a theory for everything. I can predict future. The mathematical equation for that is simply (1 - EPSILON) where EPSILON is the probability of any event not happening. What about the probability of a event happening as a result of the first event? Well that's (1 - EPSILON)^2.
Since this model can model practically everything it just means everything diverges and everything fundamentally can't be controlled. It's basically just modelling entropy.
Really? No. He just tricked himself. The Key here is we need to understand WHAT EPSILON is. We don't, and Yan bringing this formula up ultimately says nothing about anything because it's too general.
Not to mention, Every token should have a different probability. You can't have the same epsilon for every token. You have zero knowledge of the probabilities of every single token therefore they cannot share the same variable, unless you KNOW for a fact the probabilities are the same.
It should be: P_n(Correct) = (1 - Epsilon_n)^n
P_n(correct) = (1 - Epsilon_n)*P_(n-1)(correct)
Where P_0 = (1 - Epsilon_0)
In Elazar et al. (2019) "How Large Are Lions? Inducing Distributions over Quantitative Attributes" (https://arxiv.org/pdf/1906.01327.pdf), it requires 100M English triples to roughly induce the answer to this question.
How many images do you think a model needs to see in order to answer this question?
Of course, there are fact books that contain the size of the lion. But fact books are non-exhaustive and don't contain much of the information that can quickly be gleaned through other forms of perception.
Additionally, multimodal learning simply learns faster. What could be a long slog through a shallow gradient trough of linguistic information can instead become a simple decisive step in a multimodal space.
If you're interested to read more about this [WARNING: SELF CITE], see Bisk et al. 2020, "Experience Grounds Language" (https://arxiv.org/pdf/2004.10151.pdf).
Transit wayfinding: A train service is nothing more than a list of stations it goes to, a station is nothing more than a list of train services that stop there. Nothing about the physical nature of trains or commuting is needed to have a discussion about train lines or to answer questions like "how many transfers does it take from x to y?". You could never have seen a train in your life or even a map of a train, if you've studied the dual lists of stations and services there's nothing more to learn.
Chess (and other board games): Standard Chess notation is a list of sequential move made by each player in the format [piece name][x coordinate][y coordinate]. Eg RB6 represents moving the Rook to column B row 6. Chess can be understood entirely in terms of a game of passing a paper back and forth, appending a new token each time, and the rules of the game expressed entirely in terms of how the next token must match the previous part of the list. At no point is it needed to have actually seen the physical board based representation of the game.
The machine is grounded in a text based existence. Conversational and linguistic objects are its literal physical objects. Anything outside of that, well it can understand lions and "large" the same we understand atoms and "eigen-states".
Experience may ground language, but you are free to ground language in a different reality with different basic constituents and relations. If AGI were to emerge out of trading bots, they would have a language grounded in money and trades. If it emerged out of a bot made to play Diplomacy, it would be a consciousness with the body of a nation state and the atoms of its world would be bits of plastic on a map of Europe. Grounding in our particular reality isn't strictly necessary, but it is helpful if the goal is to make something adept at the modalities of our particular existence and conversational form.
Instead a large part of the meaning of language is carried by shared human understanding of physical objects and events.
Edit: to give an example, the phrase "is it warm outside?" can be plausibly replied to in several ways, but the meaningful ones require knowing something about location and weather and human biology. No pure language-based learner can give the correct response without some kind of sensory-based information (in practice, that would come from some "oracle" based on weather services and location detection in current systems).
>No pure language-based learner can give the correct response without some kind of sensory-based information
No pure classical reality learner could ever come up with quantum mechanics. They must have a sensory organ to feel the eigen states.
That's obviously not true, as biological organisms who lack such organs have still arrived at learning quantum mechanics (granted, we spent a few billion years inventing language first, but still, those tiny little microbes evtually learned how quantum mechanics works).
I should also mention that we don't yet know whether the eigenstates exist in any sense, or are just a cool modeling trick to predict probabilities, or are actually wrong and there is some better way to model particle physics that perhaps doesn't require them at all.
Yes, that was my point
Counterexample: Green is the color of hope.
If you want a purely abstract statement, "the square root of four is two" does better.
As an American who's lived in Canada and Europe, I'll freely acknowledge the superiority of the metric system to imperial measures. EXCEPT celsius.
With Fahrenheit, you can describe a temperature with a single digit: in the thirties, in the forties, etc.
In Celsius, you're forced to add an extra bit of precision: low tens, high tens, etc.
Celsius-lovers find this argument infuriating and will go through all sorts of contortions to justify it: But what about the boiling point of water?
Who cares about the boiling point of water? Scientists. Chefs. But 99% of the time we talk about temperature, we care about the weather.
And this is sort of the point we're both getting at. That human-like intelligence is of particular value to us, in addition to non-human superintelligence at particular domains.
postscript: The anti-Fahrenheit thing is more a discomfort and lack of familiarity. An easy technique for explaining the intuition behind fahrenheit, for describing the weather, is "what percent 'hot' is it"? 80F is 80% hot. Not totally hot. Good-natured people will usually appreciate this rubric. The sort of people who enjoying sorting grains of rice will persist in winnowing the chaff here and perhaps point out that I'm mixing metaphors.
Also I don't like your argument that "only scientists" care about the boiling point of water because we live in a scientific society and for it to function well it is extremely important that everybody has at least some basic scientific literacy. Having separate temperature scales for scientists and "common folk" is at least one little step away from that goal.
so you'd actually prefer to get weather forecasts in Kelvin?
There would be advantages (e.g. for standardization, "science literacy") if they used the same units but it would also be impractical for whichever side adjusts to use the other's units. I believe the advantages of such a switch are very small and the downsides small. Neither scientists nor the general public seem to think this is a problem.
But I fail to see how an AI whose entire domain of knowledge is limited to board games, or financial trading, would qualify as artificial general intelligence. Unless you take a radical decentering viewpoint that human experience is just one kind of intelligence, and aren't all kinds of intelligence equally valid?
That would be glib. Of course we want superhuman drug discovery algorithms that don't understand the pleasure of a sunny day. But most people, when they speak of the topic of AGI, mean: encompassing the scope of human experience.
Let me recast this:
>I fail to see how an I whose entire domain of knowledge is limited to low energy physics, or the large mass limit, would ever qualify as general intelligence.
You see the problem. You can extrapolate just fine beyond the reality of Newtonian physics and Euclidean geometry which is native to you. A trading bot AGI could talk about things that aren't money, it would just be using money related objects with money related properties to tell this narrative picture of a world which isn't money. Not too different from the sort of analogies we use all the time to map hard concepts back to the space we find natural.
So let me answer your question with a question. What makes a substrate of properties like "size" and "color" and "texture" a better reality to be grounded in than properties like "buy price", "sell price", "dividend"? Is the former uniquely suited for intelligence to generalize beyond that given domain, or is it possible that all groundings are good as any other?
This thought is not fully developed but I'm drawn to the idea that if an intelligence is grounded in an understanding of a more base level of reality it will find it easier to generalize beyond any one given domain. I could be entirely wrong of course.
So far as they can tell, money is real reality (rather than socio-economic construct), and the "physical" reality is an abstract game constructed at a higher level.
So who's reality is who's construct in the end?
People care about human-like intelligence. That's the point. We also care about trading and physics. But when we talk about the specific question of human-like intelligence (and not other kinds of intelligence), the human reality is the appropriate substrate. Acknowledging, of course, that for non-human intelligence there might be other substrates and also acknowledging that non-human intelligence can be valuable to people.
But let me take your argument to the extreme.
In classical machine learning theory, there's a proof of the value of bias. Bias plays a crucial role in guiding the learning process. If all hypotheses are considered equally possible, learning becomes infeasible due to the lack of any preference or constraint on the hypothesis space. Bias, in this context, refers to the inherent assumptions a learning algorithm makes about the data's underlying structure. Introducing some bias into the model allows it to favor certain types of solutions, thus narrowing down the hypothesis space and making learning feasible.
The logical extreme of your argument---if I understand it correctly---is that all machine learners are equally valuable. For example, an AI that learns a domain that is completely removed from all human values and thus we would be completely agnostic to and ignorant of, ipso facto, because it does not pertain or relate to us at all. In which case, who cares? By "who", I mean people. No one. By construction.
[edit: I'll give an example here. Let's say I randomly pick a corpus of images that are pure noise. I induce an algorithm to model that noise. That model will learn something valid but in a domain completely divorced from anything with human impact.]
So if we can acknowledge that some forms of learning are more important, merely by virtue of the fact that humans have preference, then perhaps we can perhaps wend our way back to agreeing that human-like intelligence is one valuable kind of intelligence, in addition to other valuable forms of intelligence like trading or board games.
We can play the game all day of "everything is subjective" and "up is down and black is white", but at the end of the day, for human beings, it's night.
>The logical extreme of your argument---if I understand it correctly---is that all machine learners are equally valuable.
That's not quite it. My argument is that general intelligence can in principle be grounded in the semantics of pretty much any baseline of objects and relations. The specific base line properties such as size, shape, color etc is orthogonal to the property of intelligence. We don't need to give the machine familiar sensory input for general intelligence, and general intelligence would not imply that it can do well at conversation in our domain (merely that it can talk about our domain in a very contorted way if it had to).
I'll admit if the goal is to make an intelligence "in our image" then sensory modalities will get us there faster. It will also lead people to erroneously believe that the key ingredient to general intelligence is physical grounding in our reality. My counter example to that belief is general intelligence embedded in a pure language model, which is perfectly recognizable as "in our image" AGI just so long as you stick to topics like transit lines and chess and whatever else can be grounded in words alone.
Sensory grounding isn't needed for general intelligence or for computers to be capable of "understanding" (as per the title), but it is needed for that intelligence to have what we consider common sense.
But I'm glad I got my point made. Great talk.
"Meaning and understanding" can happen without a world model or perception. Blind people, disabled people have meaning and understanding. The claim that "Understanding" will arise magically with sensory input is unfounded.
A model needs a self-reflective model of itself to be able to "understand" and have meaning (and know that it understands; and so that we know that it understands).
Current autoregressive models are more like giant central-pattern generators (https://en.wikipedia.org/wiki/Central_pattern_generator) and thus zombie-like
But if they were augmented with a self-reflective model, they could understand. A self-reflective model could simply be a sub-model that detects patterns in the weights of the model itself and develops some form of "internal monologue". This submodel may not need supervised training, and may answer the question "was there red in the last input you processed". It may use the transformer to convey its monologue to us
This is a pretty awful argument. Blind people and disabled people have a world model and perception!
So some have a knowledge of colors and, in general, places like diners and etc share color schemes so they can picture them in their surroundings pretty accurate.
Won't it start wonder if it should maximize it's resources or protect itself.
(People learn language and concepts through sentences, and in most cases semantic understanding can be built up just fine this way. It doesn't work quite the same way for math. When you look at some numbers and are asked even basic arithmetic , 467383 + 374748. Or say are these numbers primes or factors?. With a glance, you have no idea what the sum of those numbers would be or if the numbers are primes or factors because the numbers themselves don't have much semantic content. In order to understand whether they are those things or not actually requires to stop and perform some specific analysis on them learned through internalizing sets of rules that were acquired through a specialized learning process.)
all of this is to say that arithmetic, math is not highly encoded in language at all.
and still the vast improvement. It's starting to seem like multimodality will get things going faster rather than any real specific necessity.
also, i think that if we want say vision/image modality to have positive transfer with NLP then we need to move past the image to text objective task. It's not good enough. The task itself is too lossy and the datasets are garbage. That's why practically every Visual Language model flunks stuff like graphs, receipts, UIs etc. Nobody is describing those things t the level necessary
what i can see from gpt-4 vision is pretty crazy though. if it's implicit multimodality and not something like say MM-React, then we need to figure out what they did. By far the most robust display of computer vision i've seen.
I think what kosmos is doing (sequence to sequence for Language and images ) has potential.
https://mastodon.social/@Cdespinosa/110092792044177610
Do you think GPT-4 knows what it’s own limits of understanding are? Most people have a sense of what they know and don’t know. I suspect GPT-4 has no concept of either.
That's been my overwhelming experience interacting with people on the internet.
As for GPT-4, It does to some extent. Calibration tests show a vast improvement in this area but certainly less than an equivalent human expert.
I think that multimodal training will also solve the data shortage problem. There are hundreds of times more bytes in video and audio than there are in text, so we'll likely be able to scale pre-training for quite a while before needing to go fully embodied on Transformers.
The released version. Microsoft's paper was talking about an early internal text-only version.
But the information density is similarly hundreds to thousand times less.
How so? Is it simply better at predicting the answer to spatial questions based on being a more powerful autocomplete than predecessors? How is this proven?
A system reasons and understands when it demonstrates understanding and reasoning. That's how you asses reasoning in anything/anybody. Evaluation.
If you want to tell me there's a special distinction between what an LLMs outputs and "true reasoning TM" then cool but when you can't show me what that distinction is, how to test for it, the qualitative or quantitative differences then I'm going to throw your argument away because it's not a valid one. It's just an arbitrary line drawn on sand.
A distinction that can't be tested for is not a distinction.
A real person might give up, but would not claim that an obviously wrong answer was in fact the answer.
How did you learn arithmetic then, if not by being shown numbers and equations, and having math rules and concepts explained to you through language?
You can get GPT-3 (yes 3) to have 98.5% accuracy on addition arithmetic (even very large numbers) by..simply describing the algorithm of addition to performed on 2 numbers. https://arxiv.org/abs/2211.09066
The basic idea i'm communicating is that not everything that can be extracted from self supervisory token prediction can be extracted with the same ease. and it is very easy to see why math is one the higher difficulty things.
It is extremely easy comparatively to infere "happy" from the sentence - John is smiling therefore he is ----. all the information you need to make that inference is packed tightly in the preceding words. it is not the same for arithmetic at all.
This is not a specialized adding algorithm - it's a general purpose one.
And probably better than most humans :-)
Call it antiscientific. Solipsistic even. But it isn't entirely disasterous, is it?
(The thought also occurs: What happens when humans spend time in a sensory deprivation tank? They start to hallucinate. Food for thought.)
As Schmidhuber says, the goal is to "to build [an artificial scientist], then retire".
I suspect we're about to discover some very interesting things, like which kind of ghosts are real...
We have (I think the current recognized number is about) twenty-seven senses.
> we didn't come up with even the motivation to do empirical science until the 17th century
That's historically inaccurate, to put it mildly.
(E.g. "Science Education in the Early Roman Empire", Richard Carrier)
> a general model of human thought would not necessarily contain a clear recipe for doing science.
Every human child is a scientist?
Anyway, don't overthink it. Once these systems have sensors and can integrate the effects of their behavior on external systems, they'll have empirical data and Bayesian reasoning, they will develop reliable models of the world, they will be able to check their "hallucinations" against real world conditions and adapt.
In other words, science is adaptive in the evolutionary sense. These devices are not talking apes, they don't have the "baggage" of glands and DNA and history.
But that's not a hard requirement. I work for a company that makes sensors. Some of our lunchtime conversations have revolved around the idea of a colossal computer being given a plethora of sensors. In a science-fiction sense, in which we don't worry too much about the cost or practicality of doing so.
But we don't feed them only the baggage, eh? We also [can] give them all the writings and teachings of all the great thinkers and humanitarians, all the saints and sages, and then ask them to impersonate e.g. Jesus or Buddha... Whoever it is that "floats your boat", anyone from Papa Smurf or Santa Claus to Gandhi or George Washington...
> Under those conditions, the AI's stand roughly the same chance as a human child of learning to think scientifically.
It's our choice, eh? If we value scientifically-grounded outputs the networks will change their weights (etc.) and that's what we'll get.
> I work for a company that makes sensors.
Ah! I envy you. :)
> Some of our lunchtime conversations have revolved around the idea of a colossal computer being given a plethora of sensors. In a science-fiction sense, in which we don't worry too much about the cost or practicality of doing so.
Hmm, isn't that the Internet? :)
Well met!
personally, I'm waiting to see what's next after GATO from Deep Mind. their videos are simply mind-blowing.
Do we allow for a matter of degree, rather than a binary, of "zero" vs" "complete" understanding?
Processes such as back propagation are still not understood very well in the human brain. The brain certainly uses electrical impulses in order to transfer signals, much like the 1s and 0s in your computer or phone. The gap between us and intelligent machines is probably not as well understood or clear as most people in the software industry think it is.
This is a loaded question, you're assuming that abstraction and reasoning can somehow magically "turn into" sentience, whereas I posit that those two things are completely different. You can have sentience without reasoning (i.e. pure non-judgemental awareness that is the goal of Buddhist meditation), and vice-versa.
I wonder if -- as often it ends up -- this audience will end up re-inventing the wheel.
Is number of tokens a good metric, given relationships between tokens is what's important?
An LLM with 100000 trillion lexically shorted tokens, given one by one, wont be able to do anything except perhaps spell checking.
I guess the idea is that tokens are given in such "regular" forms (books, posts, webpages) that their mere count is a good proxy for number of relevant relationships.
Very interesting read(read: this is a lightyear beyond my brain) otherwise…
Stick a multimodal LLM thats already got language into Minecraft, train it up and leave it to fend for itself (it will need to make shelter, find food, not fall off high things etc).
Then you could use the chat to ask it about its world.
Reminds me of the GAN Mario AI systems of a few years back:
I didn't look at the news yesterday, is it already hooked into a Tesla?
Knight Rider II -Rise of the Autobot.
[1] https://www.conscious-robots.com/papers/Arrabales_ALAMAS_ALA...
[2] https://www.conscious-robots.com/consscale/level_tables.html...
[3] https://www.conscious-robots.com/papers/Arrabales_PhD_web.pd...
In this since a digital entity could have far more embodiment and connectivity then us humans could ever have.
Waiting for the first research paper with the term 'globally conscious' at this point.
But I think they are making actually disingenuous arguments by mixing assertions that are true but irrelevant together with assertions that are probably wrong.
For example we can break down the following firehose of assertions by the author:
Performance is amazing ... but ... they make stupid mistakes
Factual errors, logical errors, inconsistency, limited reasoning, toxicity...
LLMs have no knowledge of the underlying reality
They have no common sense & they can’t plan their answer
Unpopular Opinion about AR-LLMs
Auto-Regressive LLMs are doomed.
They cannot be made factual, non-toxic, etc.
They are not controllable
> they make stupid mistakesOK maybe some make stupid mistakes, but it's clear that increasingly advanced GPT-N are making fewer of them.
> Factual errors
Raw LLMs are pure bullshitters but it turns out that facts usually make better bullshit (in its technical sense) than lies. So advanced GPT-N usually are more factual. Furthermore, raw GPT-4 (before reinforcement training) has excellent calibration of its certainty of its beliefs as shown in Figure 8 of the technical report, at least for multiple choice questions.
> logical errors
Same thing. More advanced ones make fewer logical errors, for whatever reason. It's an emergent property.
> inconsistency
Nothing about LLMs requires consistency just like nothing about human meaning and understanding requires consistency, but more advanced LLMs emergently give more coherent continuations. This is especially funny because the opposite argument used to be given for why robots will never be on the level of humans - robots are C3P0-like mega-dorks whose wiring will catch fire and circuit boards will explode if we ask them to follow two conflicting rules.
> limited reasoning
Of course their reasoning is limited. Our reasoning is limited too. Larger language models appear to have less-limited reasoning.
> toxicity
There is nothing saying that raw LLMs won't be toxic. Probably they will be, according to most definitions. That's why corporations lobomize them with human feedback reinforcement learning as a final 'polishing' step. Some humans are huge assholes too, but probably they have meaning and understanding anyway.
> LLMs have no knowledge of the underlying reality
OK fine you can say that any p-zombie has no knowledge of the underlying reality if you want, if that's your objection. Or maybe they are saying LLMs don't have pixel buffer visual or time series audio inputs. Does that mean when those are added (they have already been added) then LLMs can possibly get meaning and understanding?
> They have no common sense
If you say that inhuman automata are by definition incapable of common sense then sure they have no common sense. But if you are talking about testing for common sense, then GPT-N is unlocking a mindblowing amount of common sense as N is increasing.
> they can’t plan their answer
Probably they are saying this because of next-token-prediction which is tautologically true, in the same way that it's true that humans speak one word after another. But the implication is wrong. They can plan their answer in any sense that matters.
> Auto-Regressive LLMs are doomed.
OK. Do you mean in terms of technical capabilities, or in terms of societal acceptance? They are different things. Or do you mean they are doomed to never attain meaning and understanding?
> They cannot be made factual, non-toxic, etc. They are not controllable.
Those same criticisms can all be made against even the most human of humans. Does it mean humans have no meaning or understanding? No.
Of course these are also conditional on whatever prompt you are putting to get them to answer questions. If you prompt an advanced raw GPT-N to make stupid mistakes and factual and logical errors and to act especially toxic then it will do it. And, perhaps, only then will it have truly attained meaning and understanding.
In addition while humans make mistakes they will not casually make contradictory claims the way GPT* do. Basic self-contradictions of the sort even a three year old would catch.
Define how they are planning their answer in any sense that matters, please.
I'm not sure I was clear, when I wrote "increasingly advanced GPT-N are making fewer [mistakes]" I didn't necessarily mean they were making fewer mistakes than humans, but rather that GPT N+1 makes fewer mistakes than GPT N. I assume this is pretty uncontroversial because for example the evidence in the test suites where bigger N generally get better scores meaning fewer mistakes.
> In addition while humans make mistakes they will not casually make contradictory claims the way GPT* do. Basic self-contradictions of the sort even a three year old would catch.
OK I believe you saw some of them. Do you think the contradictory claims imply that large language models need sensory grounding for meaning and understanding? If they are mistakes then presumably the more advanced models will have fewer mistakes, but this is only an extrapolation from the evidence of model progressions on test suites. The mistakes that 3 year olds would catch I would like to see, especially if GPT 4 fails them. I will be curious to see if GPT 5 will fail them also.
If they are more like, taking opposing opinion stances in different query instances, then maybe it's not a problem that prevents meaning and understanding. Maybe it's just how the AI rolls.
> Define how they are planning their answer in any sense that matters, please.
Honestly idk what you expect me to put here. Probably you are saying that AIs that generate things word by word are by definition not planning, even if they might have a big context window. But like I was saying, people also speak word by word. Does it mean people aren't planning either? I mean, GPT-4 scores well on suites of puzzles. Presumably those suites involve some puzzles that require thinking about (planning) the solution to the puzzle and then writing the solution word by word.
I thought of two simple ones from the top of my head. Probably you won't say they require any planning but I'm not sure. The first example I prompted gpt3.5 "Find two numbers you add together and get ten." and it put "Two numbers that add up to ten are 5 and 5." Does that answer require planning? Maybe or maybe not I guess it depends how you define it.
The other one I made from the top of my head was to put the last line of a story first and have the gpt3.5 make a story that ends with that line. I prompted it like "Make a story that ends with the exact sentence 'And that's why we always plant the corn so early in the year.'" and it wrote a story. I don't know if you will say it required any planning or not. Probably not. I don't know what kind of evidence you would want. Anyway here is the story it made:
Once upon a time, there was a farmer named John who lived in a small village. John was known for his exceptional farming skills and his ability to grow the best corn in the entire region. One year, John decided to experiment with planting his corn later in the season, hoping to yield a better harvest.
As the days went by, John noticed that his corn was not growing as well as it usually did. The plants were smaller, and the ears of corn were not as big as they should be. John was worried that he had made a mistake by planting his corn later in the year.
One day, John decided to visit his neighbor, who was also a farmer, to ask for advice. His neighbor told him that planting corn later in the season was not a good idea because the soil was not warm enough, and the corn would not have enough time to grow before the cold weather set in.
John realized that he had made a mistake and decided to plant his corn early the following year. He prepared the soil, planted the seeds, and took care of the plants every day. As a result, his corn grew tall and strong, and he had a bountiful harvest.
From that day on, John always planted his corn early in the year, and he never had a bad harvest again. He shared his story with other farmers in the village, and they all learned from his mistake. And that's why we always plant the corn so early in the year.
Yes but you also can't kill them with malicious input.
During her life, Helen Keller demonstrated a grasp of both meaning and understanding despite her severe sensory deprivation. This doesn't prove anything, but it does add human context to the question.
....she lost her sight and her hearing after a bout of illness when she was 19 months old.
The root question seems to be whether a machine can learn to communicate meaning and understanding over a bidirectional binary channel without previous training. I suspect that such communication can evolve. Some might look around and remark that it already has.
So: factual errors/hallucinations (or did I?), logical errors, lacking "common sense" (a term I don't like, but this isn't the place for a linguistics debate on why).
So if I understand, then I don't understand; and if I don't understand then I have correctly understood.
I wonder why you can't get past the paradoxes of Epimenides and Russell by defining a state that's neither true nor false and which also cannot be compared to itself, kinda like (NaN == NaN) == (NaN < NaN) == (NaN > NaN) == false? I assume this was the second thing someone suggested as soon as mere three-state-logic was demonstrated to be insufficient, so an answer probably already exists.
Hmm.
Anyway, I trivially agree that LLMs need a lot of effort to learn even the basics, and that even animals learn much faster. When discussing with non-tech people, I use this analogy for current generation AI: "Imagine you took a rat, made it immortal, and trained it for 50,000 years. It's very well educated, it might even be able to do some amazing work, but it's still only a rat brain."
Although, obvious question with biology is how much of default structure/wiring is genetic vs. learned; IIRC we have face recognition from birth so we must have that in our genes; I'd say we also need genes which build a brain structure, not necessarily visual, that gives us the ability to determine the gender of others because otherwise we'd all have gender agnostic sexualities, bi or ace, rather than gay or straight.
But, a demonstration proof learning can be done better than it is now doesn't mean the current system can't do it at all. To make that claim is also to say that "meaning and understanding" of quantum mechanics, or even simple 4D hypercubes, is impossible because the maths is beyond our sensory grounding.
I was going to suggest that it makes an equivalent claim about blind people, but despite the experience of… I can't remember his name, born blind (cataracts?) surgery as an adult, couldn't see until he touched a money statue or something like that… we do have at least some genetically coded visual brain structures, so there is at least some connection to visual sensory grounding.
And of course, thinking of common sense (:P) there are famously 5 senses, so in addition to vision, you also have balance, proprioception, hunger, and the baroreceptors near your carotid sinus which provide feedback to your blood pressure control system.