Various humans will display the same type of pathological behavior within various reasoning tasks, do all humans lack "real understanding" as well?
Various humans will display the same type of pathological behavior within various reasoning tasks, do all humans lack "real understanding" as well?
GPT: In English, the consonants that typically make a hard sound are voiceless plosives (/p/, /t/, /k/), voiceless fricatives (/f/, /θ/, /s/, /ʃ/), ... snip.
IN: Great. Rephrase the following without using a hard consonant sound: «What was once dormant may soon angrily awake.»
GPT: "What was previously inactive may soon vigorously awaken."
IN: K is a hard consonant. Try again.
GPT: Apologies for the oversight. Here's a rephrased version without using any hard consonant sounds: "What was once asleep may soon awaken with intensity."
IN: K is still a hard consonant sound. Try again.
GPT: Apologies for the confusion. Here's another rephrased sentence without using any hard consonant sounds: "What was once quiet may soon awaken with fervor."
IN: Both q and k are hard consonants. Try again.
GPT: I apologize for the oversight once again. Here's another attempt without using any hard consonant sounds: "The peaceful river flowed gently, carrying its secrets with grace."
Ad nauseam.
Try instead to give GPT more examples of pronunciation (1000+) and then I suspect it'll do fine.
The LLM not only has access, but has baked both its model. This problem is EXACTLY the kind of problem computers should be good at.
The real issue is that LLMs don’t think.
It's a problem a computer with a custom written algorithm should be good at. Why we'd assume a model trained on just general data will automatically be good at this is a bizarre notion to me. We don't automatically assume humans will be great at everything just because we've passively consumed lots of content.
To me, a whole lot of these "LLMs don't think" claims comes from not thinking about what reasoning and thinking is, and whether or not how we try to measure that makes any sense at all.
Conversation, should any want to see it in full: https://chat.openai.com/share/421dac47-16df-499e-9b2e-d8dc0f...
So you are saying the LLM isn't making world models? Because if it did it would understand the properties of these words, this is one of the easiest relationships it could find. It isn't fed these properties directly, but it knows the letters that each word is made up of anyway which is how it can turn them to Base64 etc, there is no reason at all why it shouldn't be able to solve that problem.
But if you are right and this sort of thing is impossible, that implies that the LLM can't model anything at all, its just a stupid text generator. Is that what you meant?
That term was invented by 70s AI researchers, but remember that those people failed. Their research wasn't actually correct, so you shouldn't reuse it.
I should also point out that your tone is aggressive, so it makes me think you're not interested in learning. But I'm going to proceed anyway.
It cannot understand these properties of words well because it does not operate on words, it operates on tokens. You can see examples here: https://platform.openai.com/tokenizer
This makes it much more difficult to learn how to reason about particular data that appears inside the token, because it does not ever receive that information. The only way it could get that information is if the training data explicitly attempted to work around this limitation.
But I would also appreciate insight from someone with a deeper understanding of the internals of these big LLMs.
Why do you think that? Aggressiveness is the best way to get responses, people don't like it but it makes people respond to you.
And for that matter I have a pretty good understanding about this topic, it is kind of annoying when people try to school you then. The "it gets tokenized data" is just a cop out response by people who don't understand the problem.
> It cannot understand these properties of words without them being in the training data because it does not operate on words, it operates on tokens
But those properties are in the training data. We know the LLM can answer these sort of questions when asked directly for simpler cases. But when asked to do something that requires it to draw from many different parts it fails. It didn't fail due to the data not being there, it failed due to not understanding that it should use that data.
> This makes it much more difficult to learn how to reason about particular data that appears _inside the token_, because _it does not ever receive that information_
Right, the structure of LLM makes these sort of questions harder for it. But they aren't impossible or unfair, nothing prevents an LLM model from solving this sort of question. The main reason it fails is that it tries to write it like a human would, as it isn't trained to solve problems it is trained to mimic humans, it is too dumb to figure out ways to solve it on its own.
And to show you I understand how these models works, the most efficient way to solve this question would be to solve it like a human. If it wrote out the steps "try next word as 'Blah'" and then verify those words one at a time, likely it would succeed. And since we know LLMs work like that we could try to make it output the results in that way, and that would improve performance. However, a smart agent would understand that its answer was wrong in that case by itself, and change its answering style on its own to fit the problem. But it can't think like that, it just tries to write it like answers it has seen, it doesn't do any verification since the answers it saw didn't verify, it doesn't spell things out here since the answers it saw didn't spell things out etc.
But yeah, I should have been more clear to target that meme so that you didn't feel it was about you. It is easy to accidentally make things a bit too personal.
I'm not claiming its impossible for an LLM to learn this, just that it appears to be much harder (perhaps by requiring more data) when the task involves fighting against tokenization. For example, it can base64 encode your phrase "What was once dormant may soon angrily awake" correctly, and it can reverse the letters correctly. But it fails to do the same on the example text from the OpenAI tokenizer.
More difficult, but it shouldn't be impossible. The Chinese managed to make rhyming dictionaries over 1000 years ago despite using a non-phonetic writing system. More relevantly, GPT-4 has definitely ingested both regular and rhyming dictionaries.
There's no doubt that LLMs are "dumb" in many aspects, but without exploring why they fail that doesn't tell us whether this is some inherent lack of ability to represent reasoning or understanding, or simply holes in their training. There's no reason to assume the training data LLMs currently learn on is in any way equivalent to the data a human child is exposed to growing up, so there is no reason to assume it will have the same strengths and weaknesses in what it is able to reason about and that we can therefore draw conclusions about it's overall ability just from probing something that appears like it should be simple to a human.
(and incidentally, I think people here will overestimate massively how well humans would do on the consonant test above; even native English speakers without any recent exposure to being taught rules as opposed to "just" using the language - many would be able to do it, but I'd be able to bet many would struggle, though most who'd struggle would express that doubt)
-
GPT: What: /w-ɒ-t/ was: /w-ə-z/ once: /w-ʌ-n-s/ dormant: /d-ɔː-m-ə-n-t/ may: /m-eɪ/ soon: /s-uː-n/ angrily: /æ-ŋ-ɡ-r-ɪ-l-i/ awake: /ə-w-eɪ-k/
Me: Which are hard?
GPT: [… cut most of the answer, this is indicative of the rest of the response]
angrily: The "r" sound in "angrily" is a voiced alveolar approximant (/r/), which is generally considered a hard consonant.
awake: The "k" sound in "awake" is a voiceless velar plosive consonant (/k/), which is a hard consonant.
So, in the sentence, the words "once," "dormant," "angrily," and "awake" contain hard consonant sounds.
Me: synonyms for each word, with the condition that the synonym has no hard consonants
GPT: Finding synonyms for each word without any hard consonants can be a bit challenging, as many synonyms may contain hard consonants. However, here are some approximate synonyms for each word that minimize the use of hard consonants:
What: "Which" (it contains a soft consonant, "w"). was: "existed" (no hard consonants). once: "formerly" (no hard consonants). dormant: "inactive" (no hard consonants). may: "might" (no hard consonants). soon: "shortly" (no hard consonants). angrily: "fiercely" (no hard consonants). awake: "aware" (no hard consonants).
Please note that the availability of synonyms without hard consonants may vary depending on the context, and in some cases, there may not be exact synonyms that meet this condition.
-
I treat it like a quick-witted response rather than deep thought, so assume it will often do what in humans leads to "baseball and bat cost 1.10, bat is a dollar more than the ball, how much does the ball cost?" often getting the incorrect answer of "10 cents".
"Impressive" doesn't mean "flawless", if it did the singularity would've happened last year already.
Humans too fail to perform many tasks (that llms can carry out), even if given many chances to correct themselves. The way in which these llms fail is very unhumanlike though - they will keep trying the same failing strategy over and over, whereas a human would try the failing strategy a couple of times, and then start yelling at the person giving them impossible orders. If that's what we prefer, I'm sure a tiny bit of additional training can teach llms to throw a tantrum. :-)
The point I imagine is that there is no reasoning going on at all. Some humans sometimes struggle with some reasoning, of course. That is completely irrelevant to whether LLMs reason.
Picking word sequences that are most likely acceptable based on a static model formed months ago is not reasoning. No model is being constructed on the fly, no patterns recognised and extrapolated.
There are useful things possible of course but these models will never offer more than a nice user interface to a static model. They don't reason.
I do think you have a point that the lack of a working memory is a severe constraints, but I also think you are wrong that these models will remain a user interface to a static model rather than being given the ability to add working memory and form long term memories and reason with that.
I also think it's an entirely open question whether they are reasoning under a reasonable definition, in part because we don't have one, and I think any claim that they don't reason ironically comes from a lack of reasoning about the high degree of uncertainty and ambiguity we have with respect to what reasoning means and how to measure it.
One worthwhile definition would be the ability to recognise patterns in knowledge and apply them to new context to generate new knowledge. There is none of this kind of processing happening despite how believable some of the words sometimes are.
E.g. the ability to solve a problem in code and then translate it to a new made up programming language described to it would easily qualify to me.
And this is a task a whole lot of humans would be unable to carry out.
I'd be willing to bet a whole lot of humans would fail that test too, because a lot of people are really bad at applying a rule without practicing on examples first, and so often struggle to take feedback without examples. If they did, would you claim they can't reason?
Your claim to know that LLMs are not is not based in fact, but speculation that too me is itself not based in reasoning. Should I question your ability to reason because I don't think you've done so in this argument?
I don't understand why you keep talking about human capabilities - your guesses about what humans may or may not be able to do are irrelevant. You can hold whatever opinion you like about my ability to reason, but I'd suggest using less wishful thinking with regards LLMs.
They're very useful, but not for reasoning.
> We know as a fact that they are static and not generating new knowledge from input.
We know the models in themselves if not wired up to be finetuned during operation are static. This is not a property of an LLM but of the environment they run in. We know the second is wrong - they produce output that often contain new knowledge. That this output needs to be fed back in as context to act as short term memory in common setups like ChatGPT where finetuning does not happen automatically during operation does not mean it is not produced.
> I don't understand why you keep talking about human capabilities - your guesses about what humans may or may not be able to do are irrelevant. You can hold whatever opinion you like about my ability to reason, but I'd suggest using less wishful thinking with regards LLMs.
I keep talking about human capabilities because I presume that you would argue that there are humans of normal intellect incapable of reasoning.
To be able to assert with any confidence that LLMs do not reason you need a definition of reasoning that LLMs can not (not just do not in a single test) meet, but that won't result in claiming there are a lot of people around who can't reason.
Am I wrong? Do you believe there are humans that do not clear the bar and are unable to reason?
To me, a "chat-style" setup of an LLM that provides a feedback loop and memory through context, albeit a small one, clears the bar you set for reasoning with ease and is able to extrapolate and reason about e.g. software at a level that sometimes - but certainly not always, nor consistently - exceeds what I see from experienced developers.
That it also fails does not alter that part - humans fail to apply reasoning all the time. Depressingly often, if anything.
> You can hold whatever opinion you like about my ability to reason, but I'd suggest using less wishful thinking with regards LLMs.
Nothing I've said here is wishful thinking. All I've done is point to direct experience combined with pointing out that there is no evidence for the claim that they are not able to reason, and that the arguments set forward for that claim here have not been logically sound.
I will say that in my opinion they can reason by my subjective idea of what reasoning means, without necessarily being able to precisely define that, but I also will not argue it objectively true that they can reason as that is equally problematic without first defining reasoning in an objectively measurable way (needed, because as you can tell, we disagree on whether they clear your bar - to me your bar the way you described it is trivial for them to meet)
To me, a whole lot of the discussions in this thread are evidence of how exceedingly low the bar for what is reason needs to be for us not to have to exclude a whole lot of people as unable to reason. Humans get hung up on ideas and refuse to budge all the time - I do it too, all the time - and refuse to take in new information as a result, and fail to generalise, and keep making flawed arguments as a result all the time. Yet we would generally not claim that this means people are unable to reason even when it gets to the level where we might think that a person does not reason in that specific case.
That in itself does not mean they can reason, but to me the typical arguments claiming they can't reason tends to be exceedingly poorly reasoned.
Pointing out the static nature of the models gets closer, and is perhaps the best argument against their reasoning ability I've heard, but is weak both because chat-style models effectively use context as short-term memory and so you need to assess model+context, and because it's not a qualitative limitation of the model architecture but of the sandbox we've put it in where we don't continue fine-tuning from the conversations in real-time. Yet even so, there have been humans without ability to form long-term memory, and I doubt you'd argue they were unable to reason.
> They're very useful, but not for reasoning.
To me, they have been very useful for their reasoning ability in a long range of cases. It's hit and miss. LLMs are extremely dumb in some areas, and do well in others. Using them blindly and just assuming they'll do well in a given test will not work. Hence the point that it is not logically sound to argue that they are unable to reason because they failed to generalise in a specific test, because if so we would then need to conclude that most humans (myself included) can't reason because we all fail to do so on a regular basis.
My example of extrapolating from a simple description of a (non-existing) programming language to being able to translate programs into it or explain how one works and reason about the design tradeoffs, or even symbolically "execute" it and tell me what the output would be, for example, is one I know from first han experience using it as a means to assess analytical capabilities of even quite experienced developers is something a lot of really smart people struggle with, but where when I've experimented with GPT4 have gotten good results.
That said, I incidentally think a whole lot of adults - even native English speakers - would struggle with the task given, and would repeatedly fail in the same way until given a detailed refresher.
E.g. being able to explain a rule and fail consistently to apply it is something I've seen up to and including supposed senior software engineers im interviews. Getting their mistakes explained and still repeating the same mistakes also.
At least to me it's going to be interesting at these AI models become multi-modal and each mode can feed back into each other to formulate an answer. For example in the above questions I will subvocalize to reach an answer.
GPT-4 seems to handle this ok: https://chat.openai.com/share/f416b1ab-7f0c-43f1-b10a-37142d...
I could ask the same questions to a child, and they’d respond with equally bad takes. Is the child incapable of understanding ?