I read the prompt, and I expected that this was the beginning of some kind of fiction. In my mind, it sounded like I was reading the beginning of a somebody’s dream. What does it even mean to understand something? Because naively, it looks very much like GPT-3 and I have a shared understanding of the first prompt.
Do I actually think the model understands like a human does? No. But I would bet that, in isolation, the part of my brain which processes and generates language might not understand much either...
Or maybe I’m a bot and neither I nor GPT-3 understand anything at all. Beep boop
Yes it does. A model that latches on superficial frequentist links between words is much less impressive than one that would understand what those words actually mean, and the latter is how most humans use words. The former is just chinese-rooming, the latter is understanding. Of course, a model that is chinese-rooming something like a coherent text is impressive, but it is less impressive than one that would demonstrate actual grasp of the fact that words mean something.
If the claims about GPT-3 were accurate, there'd be a lot less of a flare-up about it. Don't claim your software does what it can't.
But it does, that is the point the author is making - GPT cannot understand anything. It's a very silly argument to try and reduce your linguistic perception to be on the same level as GPT just to try and show that it possibly is understanding the text. The author's queries do a very good job at demonstrating that GPT at the very least cannot even tell when it is being asked a question. As you say it isn't unexpected, GPT is a statistical fitter over a selection of language features trained over a large dataset, it isn't intended to perform well in these scenarios - which is exactly the author's point. In order to perform well it would have to have some capability of understanding the sentence. I think it's more apt to say GPT is only capable at recognizing sentence features, it can't understand anything, at most it has built up a vague relationship between certain sentences and possible continuations.
It’s a distracting anthropomorphism to even attempt ascribing “understanding” to a model like GPT-3. An assessment of its useful capabilities should be through an honest effort to get it to do something - and should of course include consideration of the effort/intelligence required to do so. Marcus knows enough to know this set up is inappropriate, so the article reads as disingenuous.
I wonder if you can force "an attempt" at answering a question (to flagrantly anthropomorphize) by following the question with something like "The answer is ..."
This is only true if we assume GPT was never trained on satire or intentionally absurd text. But there's no reason to think this. Because it continues a bad prompt in an absurd or comical way does not demonstrate it doesn't "understand" common facts. If you treat GPT as a conversation bot and expect it to call you out when you give it an absurd prompt, then it is your expectations that are wrong.
Well then you can justify it outputting anything at all.
We are not talking about nonsense, we are talking about unexpected input, which will happen in any real-world situation. It might be text about something a kid, or depressed person has done. The text about stirring with a cigar is firmly withing the realm of plausibility.
And what happens is that GPT seemingly goes off the rails.
Furthermore, I do not think we can assume that if it were trained on examples of satire or intentionally absurd text, it would perform better on such prompts - in fact, I would not be surprised if its performance would deteriorate on many prompts, both straightforward and tricky ones, if given such training.
Now I am wondering if you need a theory of mind before you can begin to understand satire...
There can be no such thing as a bad prompt from gpt3's perspective. A bad prompt is one where the user has a specific purpose which is not expressed. It's bad because you know beforehand that gpt3 can not align with it.
Someone pours grape juice into a bottle and becomes worried that is not safe to drink. GPT3 correctly grasps that there is a hidden context to this weird prompt, however when given no other information it guesses that this hidden context is something known only to the hypothetical character in the prompt. I would probably do this too.
When you give it the correct context (this weird prompt is a logic test) then it gives you the answer you expected.
That is interesting. How does it indicate its understanding that there is a hidden context?
"Responding to a customer inquiry with 'noone cares, go away' wasn't really a failure on the part of the model. Rather, the model was simply creating a performance-art piece commenting on the way capitalism drives an emotional wedge between 'providers' and 'consumers'. Try fine-tuning on some economics journals to get that out of its system."
In any case a language model is a language model. It has no other ability than calculating the probabilities of sequences of tokens. Leaving aside the question of how something like that can have "real understanding" just by being embodied in the world, how do you even "embody" a language model? I'm genuinely curious to hear how far you have thought about that and how clearly.
I mean in practical terms- you have a trained language model. You have a robotic body (not your own). How do you put them together to produce an embodied agent? What are the intermediarey steps that lead to a robot that can use its language model to... (what does an embodied language model do)?
What is meaning if not illusion?
We don't _really_ know what the physical manifestations of meaning and form are in the brain... they're just concepts we invented.
If anything, GPT-3 is suggesting that either:
1. Tasks which were previously thought to require meaning actually turn out only to require form.
2. Meaning and form are more related than previously thought.
Both are interesting findings imo, but 2 would be huge, especially if it suggests how the brain might work. Could meaning be an emergent phenomenon of form?
If you look too deeply it quickly gets philosophical.
That's exactly what we're doing. And we're never given the "answer sheet" to figure out whether we understood the platonic, capital T Truth, or whether we just learned a spurious correlation. We just keep getting more of those inputs. Which is why it seems to me that an unsupervised sequence prediction model like GPT-3 is the only sort that could ever give rise to something akin to human consciousness.
The big differentiator seems to be that with a pure text sequence model, inputs go in, but the outputs don't have any control over future inputs. It isn't structured to have anything like agency, just passive observation and prediction. But a useful "understanding" in a human sense is related to what can be done with that understanding to enact change in the environment. I don't know how you would teach it that without giving it a Reddit account and setting it loose.
> But then I (as always I can only speak for myself, everyone else could be a p-zombie for all I know!) also have a qualitative experience of trees, and words, and an experience of meaning and understanding. If you look too deeply it quickly gets philosophical.
I'm not so sure I have those things. I'm glad you do. That's one reason I'm never going to do ketamine.
At the end when you look long enough it seems to call into question the very nature of consciousness.
To do a "farduddle" means to jump up and down really fast. An example of a sentence that uses the word farduddle is: One day when I was playing tag with my little sister, she got really excited and she started doing these crazy farduddles.
According to my understanding of the concept it must know something about meaning and is able to reason about it if it was able to generate this.
- Um, it actually already is a word. Tnetennba.
- Good heavens, really? Could you, uh, use it in a sentence for us?
- "Good morning. That's a nice Tnetennba" [1].
______________
[1] Moss from IT Crowd on Countdown:
https://youtu.be/g9ixvD0_CmM?t=52
Edit: to clarify, if you don't know what a word means, just seeing it used in a sentence won't necessarily tell you much about its meaning, so that a language model was able to generate a phrase with the word in it doesn't necessarily tells us it understands the word's meaning.
In any case, it's a language model. It has no ability to "understand" anything. It can compute the probability of a token to follow from a sequence of tokens, and that's all. There's no "understanding" there, nobody made it to understand anything.
Edit: btw, "learning" in the context of machine learning is more of a term of trade with well-established connotations. For example, we have Tom Mitchell's definition of a machine learning system as "a system that improves its performance over time", etc. We don't have similarly established definitions for the "understanding" terminology. Hence my request for clarification. I literally don't understand what you mean that GPT-3 "understands" metaphorically.
I think the more interesting question is the definition of meaning. I am thinking about meaning here as the relationship between symbols. So if you can explain what a words means, you can give a definition in terms of other words. If you "understood" what a word means, you have not just memorized the definition but can apply the word in unseen contexts.
>> So if you can explain what a words means, you can give a definition in terms of other words.
Suppose I give you the following mapping between symbols: a -> p, c -> r, d -> k, e -> j.
Now suppose I give you the phrase: "a a a c a d e e a c"
I gave you a definition of each symbol in the phrase in terms of other symbols. What does the phrase mean? Alternatively, what do the symbols, themselves, mean?
Obviously, you can't say. Being able to give the definition of a word in terms of other words presupposes you understand the other words, also. So, just because a language model is using a word doesn't mean it knows its meaning- only that it uses the word.
However, You seem to be making the Chinese Room argument. If you define meaning such that either no computer program could possibly "understand" meaning or it is unverifiable if it does, I don't think it makes much sense to have a discussion if GPT-3 does. Is there a test that a model could pass that would convince you that it "knows" meaning according to your definition?
My comment is relevant to the question of whether GPT-3 has "understanding" or not, because in order for GPT-3 to understand the meaning of a word A in terms of a meaning of a word B, it needs to already know the meaning of the word B. However, this is what we wish to know, whether GPT-3 knows the meaning of any word. Observing that GPT-3 can use a new word in the place of a different word doesn't tell us whether it knows the meaning of the original word.
As of yet, no, there is no formal test that would convince me or a majority of reserachers in AI that a model "knows", "understands" or anything like that. The reason is not that I am too stubborn, say. Rather there simply aren't such tests available yet. One reason for that is that we don't, well, understand what it means to "understand". We don't have a commonly accepted formal definition of such ability. Without that, we can't really design tests to prove that some system has it.
The take away is that it will be a long time before we can know for sure that a system is displaying intelligence, understanding, etc. This may be unsatisfying- but the alternative is to design meaningless tests that prove not what we are trying to prove and proclaim the goal proven if the tests pass. This does not go well with the purpose of scientific endeavour, which is to acquire knowledge- not pass tests and make big proclamations about winning this or that competition.
In short, I'm not saying that computers can't have understanding, or that we can't know if they do. I'm saying that right now, these things are not possible, with current technology.
It correctly interprets the first part of the prompt as 'farduddle ~= jump' and the second part as an instruction to generate a sentence containing farduddle, possibly utilising a corpus of existing sentences containing jump in the context of 'really fast'. But that's also a series of instructions you could imagine as a DSL a relatively simple program could parse and generate a satisfactory response to. Which I believe the OP is classing as 'structure' since it's just performing translations based on familiar syntax. Understanding the concept of 'jumping' is a step further, before we get into the more philosophical stuff about qualia and whether things that can't jump can ever truly understand the experience of jumping...
If the parent comment was making some kind of Chinese Room argument than I think that's not very helpful for the discussion. "GPT-3 learned nothing about meaning because no computer program can do that by definition".
But I don't think they were trying to make that kind of argument as you seem to be suggesting because they said they were unsure if it learned anything about meaning.
We don't really know what "understanding of the world" means in humans. We just "see it when it's there".
We might be radically different from GPT-3, or we might not. Our way of learning is different in any way.
Something that came to my mind: Various GPT-3 answers resemble answers given by children: Mostly correct, but having misunderstood some crucial point.
In real human learning and conversation these points are easily corrected by feedback by explanation: "You see, the point is no one wears bathing suites to work".
Which would then be incorporated as new wisdom.
Maybe this feedback-mechanism is what GPT-3 is missing. Maybe we should talk to it.