-
GPT: What: /w-ɒ-t/ was: /w-ə-z/ once: /w-ʌ-n-s/ dormant: /d-ɔː-m-ə-n-t/ may: /m-eɪ/ soon: /s-uː-n/ angrily: /æ-ŋ-ɡ-r-ɪ-l-i/ awake: /ə-w-eɪ-k/
Me: Which are hard?
GPT: [… cut most of the answer, this is indicative of the rest of the response]
angrily: The "r" sound in "angrily" is a voiced alveolar approximant (/r/), which is generally considered a hard consonant.
awake: The "k" sound in "awake" is a voiceless velar plosive consonant (/k/), which is a hard consonant.
So, in the sentence, the words "once," "dormant," "angrily," and "awake" contain hard consonant sounds.
Me: synonyms for each word, with the condition that the synonym has no hard consonants
GPT: Finding synonyms for each word without any hard consonants can be a bit challenging, as many synonyms may contain hard consonants. However, here are some approximate synonyms for each word that minimize the use of hard consonants:
What: "Which" (it contains a soft consonant, "w"). was: "existed" (no hard consonants). once: "formerly" (no hard consonants). dormant: "inactive" (no hard consonants). may: "might" (no hard consonants). soon: "shortly" (no hard consonants). angrily: "fiercely" (no hard consonants). awake: "aware" (no hard consonants).
Please note that the availability of synonyms without hard consonants may vary depending on the context, and in some cases, there may not be exact synonyms that meet this condition.
-
I treat it like a quick-witted response rather than deep thought, so assume it will often do what in humans leads to "baseball and bat cost 1.10, bat is a dollar more than the ball, how much does the ball cost?" often getting the incorrect answer of "10 cents".
"Impressive" doesn't mean "flawless", if it did the singularity would've happened last year already.
So you are saying the LLM isn't making world models? Because if it did it would understand the properties of these words, this is one of the easiest relationships it could find. It isn't fed these properties directly, but it knows the letters that each word is made up of anyway which is how it can turn them to Base64 etc, there is no reason at all why it shouldn't be able to solve that problem.
But if you are right and this sort of thing is impossible, that implies that the LLM can't model anything at all, its just a stupid text generator. Is that what you meant?
That term was invented by 70s AI researchers, but remember that those people failed. Their research wasn't actually correct, so you shouldn't reuse it.
I should also point out that your tone is aggressive, so it makes me think you're not interested in learning. But I'm going to proceed anyway.
It cannot understand these properties of words well because it does not operate on words, it operates on tokens. You can see examples here: https://platform.openai.com/tokenizer
This makes it much more difficult to learn how to reason about particular data that appears inside the token, because it does not ever receive that information. The only way it could get that information is if the training data explicitly attempted to work around this limitation.
But I would also appreciate insight from someone with a deeper understanding of the internals of these big LLMs.
Why do you think that? Aggressiveness is the best way to get responses, people don't like it but it makes people respond to you.
And for that matter I have a pretty good understanding about this topic, it is kind of annoying when people try to school you then. The "it gets tokenized data" is just a cop out response by people who don't understand the problem.
> It cannot understand these properties of words without them being in the training data because it does not operate on words, it operates on tokens
But those properties are in the training data. We know the LLM can answer these sort of questions when asked directly for simpler cases. But when asked to do something that requires it to draw from many different parts it fails. It didn't fail due to the data not being there, it failed due to not understanding that it should use that data.
> This makes it much more difficult to learn how to reason about particular data that appears _inside the token_, because _it does not ever receive that information_
Right, the structure of LLM makes these sort of questions harder for it. But they aren't impossible or unfair, nothing prevents an LLM model from solving this sort of question. The main reason it fails is that it tries to write it like a human would, as it isn't trained to solve problems it is trained to mimic humans, it is too dumb to figure out ways to solve it on its own.
And to show you I understand how these models works, the most efficient way to solve this question would be to solve it like a human. If it wrote out the steps "try next word as 'Blah'" and then verify those words one at a time, likely it would succeed. And since we know LLMs work like that we could try to make it output the results in that way, and that would improve performance. However, a smart agent would understand that its answer was wrong in that case by itself, and change its answering style on its own to fit the problem. But it can't think like that, it just tries to write it like answers it has seen, it doesn't do any verification since the answers it saw didn't verify, it doesn't spell things out here since the answers it saw didn't spell things out etc.
But yeah, I should have been more clear to target that meme so that you didn't feel it was about you. It is easy to accidentally make things a bit too personal.
I'm not claiming its impossible for an LLM to learn this, just that it appears to be much harder (perhaps by requiring more data) when the task involves fighting against tokenization. For example, it can base64 encode your phrase "What was once dormant may soon angrily awake" correctly, and it can reverse the letters correctly. But it fails to do the same on the example text from the OpenAI tokenizer.
More difficult, but it shouldn't be impossible. The Chinese managed to make rhyming dictionaries over 1000 years ago despite using a non-phonetic writing system. More relevantly, GPT-4 has definitely ingested both regular and rhyming dictionaries.
There's no doubt that LLMs are "dumb" in many aspects, but without exploring why they fail that doesn't tell us whether this is some inherent lack of ability to represent reasoning or understanding, or simply holes in their training. There's no reason to assume the training data LLMs currently learn on is in any way equivalent to the data a human child is exposed to growing up, so there is no reason to assume it will have the same strengths and weaknesses in what it is able to reason about and that we can therefore draw conclusions about it's overall ability just from probing something that appears like it should be simple to a human.
(and incidentally, I think people here will overestimate massively how well humans would do on the consonant test above; even native English speakers without any recent exposure to being taught rules as opposed to "just" using the language - many would be able to do it, but I'd be able to bet many would struggle, though most who'd struggle would express that doubt)