Seriously just watch. He's not actually going to be able to coherently define his "reasoning" in a way that can be tested.
Google gives the following definition of the verb "reason":
> think, understand, and form judgments by a process of logic.
LLMs do not think, they do not understand, and they do not form judgments. They do not come to their own conclusions. They do not have the physical capability. They are statistical models, nothing more.
> LLMs obviously reason and understand by any evaluation that can be carried out.
Uh-huh. Sure, Jan.
"LLMs don't reason because they don't understand" is not the bastion of genius you think it is. It's a circular argument that relies on whatever bespoke interpretations you have cooked up.
They don't form judgement or conclusions? Sure looks like they do. So what's the difference ?
What is GPT-4 doing then when it correctly looks like it is reasoning and what's the difference between that and "real" understanding or reasoning.
Such a huge difference I should be able to test for it. Don't understand how you can tell me what I'm seeing isn't real reasoning but fail to provide a way to empirically determine the difference.
>Uh-huh. Sure, Jan.
Yeah. https://arxiv.org/abs/2212.09196 Many other evaluations to carry out.
Instead, could you operationally define "reason" in a way that a human is, say, 90 % likely to pass the test and GPT is 10 % likely to do?
The problems are presented in a way that make it difficult to solve. The vision problems will require the equivalent of an artificial visual cortex, something we are seriously lacking in artificial intelligence at the moment. Image to text won't cut it here.
For the text there could be tokenizer issues. LLMs don't really have any problem with abstract analogical reasoning https://arxiv.org/abs/2212.09196
Another issue is the vision side. The vast majority of multimodal models are working on essentially an image to text objective task. That won't cut it here. We need the equivalent of an artificial visual cortex. We don't have that yet
If you're saying that we can't use a problem if any analog of that problem has ever been described, you seem to be arguing more strongly that it is a general intelligence than I am.
That’s not intelligence that’s computers having better memory than humans. Useful, certainly, but hardly skynet.
If you're saying that it is not possible to change the details enough to avoid the model being able to answer that type of question, I think you are admitting that the model has learned a generalized ability to answer questions of that class, and is not actually using its memory to answer at all.
I don't care about whether it learned that generalized ability from seeing examples of the question and answer, which it then deduced an algorithm for and generalized -- that's how most people learn most things.
As an aside, I’m really starting to hate these threads on here, people are constantly reading words that aren’t there in search of gotcha-it’s-skynet. It’s not. It’s just pattern matching and randomness with a giant amount of information encoded.
I'm not sure what to say, other than that if you'd like to have less frustrating conversations, you could do better than showing up with hearsay where someone asked a question they thought was unique, but it wasn't, and it can't be modified to be unique and then asked again, and you aren't willing to tell us what it was, and possibly don't know yourself.
It is not possible to have a serious conversation about your claim, and that's not because it is being intentionally misunderstood.
> skynet
You're the only person mentioning skynet. The conversation is about a ridiculous claim made up-thread that GPT-4 cannot reason or understand anything, which is disprovable within a few minutes of using it thoughtfully.
The vision problems will require something much more than an image to text objective task. It will require the equivalent of an artificial visual cortex. We don't have that yet.
For abstract analogical reasoning, LLMs don't have a problem with that. https://arxiv.org/abs/2212.09196
First the vision problems will require the equivalent of an artificial visual cortex, something we are seriously lacking in artificial intelligence at the moment. Image to text won't cut it here.
For the text, LLMs don't really have any problem with analogical reasoning https://arxiv.org/abs/2212.09196
Complex behavior arises from simple systems all the time. You can't prove that these systems don't reason, no matter how loudly thou doth protest.
We are talking about a model of language and our expectations for what the words would look like if a person were using language to reason with. People are confusing one thing (their own interpretation of what they are reading) with another (a statistical model that is being driven to align with certain specific human expectations).
The alternative is having 1000 different results for different kinds of arrows, and averaging out the results for the ones similar to the input arrow.
An LLM is working on text tokens, it’s trying to give the most statistically common next token based on everything that’s been fed into it. Does that statistical model abstract the objects and concepts it talks about? Eh? I don’t know
vvv
But, coming at this as *respectfully and curiously* as I can here:
^^^
I guess I’m a little more skeptical about attributing magic to something that I know is only working from words as a source.
Like if it’s seen NOUN really close to VERB a lot, it can “assume” (read: bias output) that NOUN will VERB in the context it sees it in, that’s an abstraction! But a weak one right? Language is pretty inaccurate. The fact that NLP was a whole booming research field a few years ago tells me that just parsing language consistently is a difficult problem.
So, I know I can’t give you any convincing argument there since we’re already mentioning how people are giving untestable definitions of “reasoning” and “abstraction” in this thread. And im no better in that respect here, BUT I guess my best inkling of what feels “off” is a lack of precision, inherent to the medium in which it operates which makes it feel like a true abstraction is impossible in an LLM.
Your intuition about the architecture might be causing the distrust. Yes, the ends of the model are words in and words out (although GPT-4 is now multi-modal and can accept images as input too), but that doesn't mean the calculations in the middle are dealing in words. The "language" in "large language model" is misleading. It's a large generalized model that imports and exports language.
It all still feels rather “attribute magic to complexity” to say that there’s some human element skimmed off the patterns of our language, hidden in the coefficients somewhere.
I’m also not a magician, so really who am I to say :)
However, that feels a bit magical to me at the moment. Still interesting to be along for the ride humanity is going on!