GPT-4 performs better at Theory of Mind tests than actual humans
twitter.com
twitter.com
So I would actually be a bit surprised if the models failed this type of test regardless of whether they have any understanding of theory of mind.
Note that I’m making no claim about how intelligent or understanding of humans or neurotypical these models are. My claim is that it’s very hard to tell given how incredibly knowledgeable the models are.
But these same devices, running the right software and given the right parameters, do something that seems impressively close to intelligence. And I suspect that my brain, with its software somehow removed, would not seem very intelligent.
On the other hand, a human baby will spontaneously generate language and intelligence, and I haven’t heard of a language model doing that. (The experiment has, sadly, been done. https://en.m.wikipedia.org/wiki/Nicaraguan_Sign_Language )
Philosophy often asks questions that are not yet reached by current technology but will be soon enough.
When a model is given insufficient context beyond the question, it may generate responses based on its best guess. This situation can be compared to abruptly waking someone up in the middle of the night and demanding an immediate response to a question.
In contrast, when humans are asked to answer questions in a test setting, they are aware of the larger context and the importance of providing accurate answers.
"Autocompletion" isn't an adequate explanation when neither the prompt nor the output has ever existed before. Something else is going on... something that obviously has a lot in common with how our own minds work.
That's not a lot of different sentences when you consider the combinatorical explosion.
I could also understand an argument saying that it's impossible for a computer to think, but I don't think I'd agree with it.
I have seen many claims that transformer models can't "think" but I have not seen any purposed tangible tests to verify or refute the claim.
Simpler models look very good on the surface, but there are some examples where this starts to break own. Example:
"Is it legal for a man to marry his widow's sister?"
The catch here is the man is dead, which is the reason he can't marry. GPT-4 gets this right. Simpler models however can be prompted to perform step-by-step reasoning where they figure out that dead men can't marry, and that a "man's widow" means he's dead, but the model still can't be coaxed into figuring out that he can't marry in the scenario above.
I'm guessing GPT-4 still has similar hangups where it gets stuck and can't get out no matter the coaxing, although I don't have any examples handy. Anyone?
This doesn't mean GPT-4 is a stochastic parrot, just that there are still flaws in its reasoning abilities.
Turns out posthumous marriages are legal in some places so a French man can legally marry his widow’s sister assuming it doesn’t run afoul of other laws like (which I didn’t check the legality of) polygamy.
LLMs are great at being interfaces to the data they were trained on. Ask them anything beyond that, and it becomes clear they're not capable of thinking like humans do, i.e. producing novel ideas.
It's kind of absurd we're having this discussion, TBH. The only reason it's difficult to distinguish mimicry from thinking is because of the sheer volume of data we're talking about here, and some clever algorithms that find correlations in it, and can synthesize a human-like response.
This is exacerbated by the fact that human education is built around memorization, and that we equate intelligence with being able to regurgitate previously known facts. So if this is your measuring stick, then by all means LLMs can be considered intelligent agents capable of rational thought, when in reality this is far from the truth.
>the sheer volume of data we're talking about here, and some clever algorithms that find correlations in it, and can synthesize a human-like response.
If you swap the volume of data from "terabytes of UTF-8" to "decades of continuous input from the 5 senses," this doesn't sound very different than what you or I do already. The fact that it's difficult to differentiate mimicking from true thinking strongly suggests that there really isn't a difference, or at minimum not one that has any real world practical significance. If there is a scenario where the difference does have practical significance, what scenario is that? Because I sure can't think of one.
If this so-called “fake” thinking is indistinguishable from “real” thinking, is there really a difference?
On the contrary, isn't "step by step" reasoning exactly what we tell students to do to become better at solving complicated problems?
Especially when you arrive at this imaginary distinction with "It's fake because it has to be fake"
If thinking/understanding/reasoning/whatever can be fake, it should be testable. It should be a conclusion that can be reached from results not baseless speculation. I mean, what kind of huge difference can't be tested for ?
If you tell me "this is fake gold" then there are numerous ways to distinguish fake gold from it's real counterpart, mostly by testing its physical properties.
If the results and properties from "fake" [insert] and "real" [insert] can't be distinguished then well you've just made up a distinction that doesn't meaningfully exist.
I agree with everything else you said, but this is a point I disagree with.
Just because something isn't testable, doesn't mean it doesn't exist. Theoretically, the state of facts could be such that 1) human minds can think (as in, have consciousness and perform the process that is known to us as "human thinking") and 2) machines cannot think (as in, they don't have consciousness, but are literally just mechanisms that appear to provide same output as humans for a certain set of inputs).
I don't think it's true, in fact, I don't think consciousness itself is anything more than a convenient illusion, but I still must point out that "what is real must be measurable" is not a valid argument.
A more precise formulation might be "if thinking can be fake, the only way we can prove it is if it's testable - otherwise it doesn't make sense to discuss it".
If it's not distinguishable then there's no difference period. Of course that won't stop people from making up differences.
An argument stemming from "well humans could have this special sauce that machines can't have...just because", is not something to take seriously.
It does not mean they are (or work) the same way.
That's true only if you believe that the world only exists inside humans' minds.
If you believe in objective reality, then things that are different but not distinguishable by humans can exist.
In other words, beyond a certain point it's impossible to tell.
At the same time, it's just as impossible for me to know that another human is sentient, even if we meet face to face.
(So it's pretty much a moot point.)
The great thing for cognitice scientists /linguists is that we now have a quantitative, precise framework and no longer need to talk in terms of the folk intelligence science of the past.
The fundamental concept behind LLMs is to allow the model to autonomously deduce concepts, rather than explicitly engineering solutions into the system.
The self-reflection part is probably true, but that’s not strictly necessary to understand concepts
> what concepts are
How do you know concepts aren’t just statistics?
You're essentially requesting the goalpost be moved.
The cognitive tests we rely on (turing test, chinese room etc etc ) are woefully outdated and inadequate for our time
The goalposts will always be moved btw, because our experience of intelligence is subjective and we 'll never have an objective measure of it. At some point we will stop moving them because we ve ran out of ideas. At that point we can say we have a facsimile of our intelligence
However:
> but clearly dont have a model of the world, self, or others because they have not been engineered in
Neither did we.
> The hypothesis that, by scaling the giant clockwork, these things will magically emerge is .. magical and unproven.
Our sapience was and is an emergent phenomenon, (superstition aside) was that magic?
Wouldn't that be a table of sentences less than 7 words long?
100^7 = (10^2)^7 = 10^14
EDIT: 0. How many times did it take this test? 1. Where's the code and reproducible results? 2. Okay, give it a completely new test. How about that? 3. Who did it cheat off of?