You can argue it's just an imitation of "understanding" and not the "real thing", but how good does an imitation have to be before it's functionally identical to the real thing?
You can argue different cherry-picked examples seem to demonstrate a lack of understanding, but does that mean it doesn't possess understanding in general or just that it doesn't understand those particular topics?
I agree GPT-4 is not a human-like AGI, and that the people expecting it to behave as one or expecting a future version right around the corner that does are likely in for disappointment. But at the same time, LLMs are capable of writing coherent, never-before-seen code from natural language instructions, solving basic never-before-seen natural language logic puzzles, and playing half-decent chess moves from never-before-seen positions[1], all in a single model that was never specifically trained to do any of those things. To say those tasks don't require at least some level of intelligence or "understanding" seems far fetched to me.
To be clear, an example does not necessitate cherry-picking. Cherry picking specifically requires one to ignore evidence to the contrary. There's a certain irony here. The reason I state this is because the criteria upon which you give me is impossible. The list of examples is non-exhaustive. Nor do I have infinite time or space in which to provide these examples. The example was shown because it demonstrates a simple question that we'd expect any reasonable understanding creature, that knows what a pound and kilogram are, to be able to answer. The LLM demonstrates that it has the requisite knowledge, that it knows the relation between the two units, but it fails to make the correct conclusions which requires the understanding part: putting the knowledge together. The follow-up question also is an example, and a different one for that matter.
I'm not sure your link creates a compelling argument. This is despite the fact that Chess itself is not generally considered a good setting for testing understanding. I mean we've pretty much agreed on that even before Deep Blue.
We need to be very careful in how we state things and try to interpret others. I hope I have not mischaracterized your claims. But I'd also encourage you to not be so antagonistic with others, especially while demonstrating the very thing you accuse them of. Internet comments are not academic and that's okay, but we should still try to be friendly. Not that a little poking won't happen.
I guess it's fair to say my examples are cherry-picked too (though I didn't really give specific examples in my comment so much as entire general categories of problems ChatGPT is known to be proficient at solving). But they aren't so cherry-picked as for "random chance" or "that example was in the training data verbatim" to be possible explanations. It's not like I had ChatGPT answer 100 billion questions and am only showing you the top 0.1%. It's very common for ChatGPT to be perfectly correct even when answering logic puzzles or coding problems not in its training set. So if not those then what's your explanation other than "understanding"?
On the flip side, I don't think it's possible to prove ChatGPT does not posses understanding with individual examples. (And it seems you agree?) Those are easily dismissable as just "ChatGPT isn't great at that particular task". Particularly given how many other tasks ChatGPT is great at.
Scott Alexander figured this out back in 2019 before GPT-3 was even a thing (let alone ChatGPT), I think it's worth a read: https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-g...
> The example was shown because it demonstrates a simple question that we'd expect any reasonable understanding creature, that knows what a pound and kilogram are, to be able to answer.
Not really though. The "pound of bricks or pound of feathers" question is specifically designed to trip up humans, and the specific formulation you used seems specifically designed to trip up LLMs (by playing off their tenancy to pattern match common sayings), yet despite those disadvantages GPT-4 succeeds.
Further, I don't think "I'd expect even a child to be able to answer this" is a good metric. ChatGPT isn't human, so we shouldn't expect it to be good at everything humans are good at just because it's good at some things humans are good at. Again, I'm not trying to argue ChatGPT is a human-like AGI.
It was where you said cherry picking. I'll admit it was an over correction given your follow-up, but internet comments (especially when disagreement exists) end up being combative. Many times unintentionally.
The point of the specific example is not that it is a trick question. It is more that the answer is self contradictory. That's the key part and what demonstrates that it doesn't understand. The second example is that follow-up questions do not result in a convergence. This is related to cognition but not as strong.
A bias to pattern match is an issue but you're right that it doesn't disprove sentience or even understanding. But a self contradiction does demonstrate a lack of understanding. If a human gave that exact answer, you would be very confused and how someone could be so stupid. Remember, it does currently identify the relationship between pounds and kilograms, fully explaining even the true base comparison through common units. But it still gets it wrong. That's the critical part. That's overfitting to the statistical nature and that this takes far more priority than meaning. Getting it wrong is one thing. Getting it wrong and using the right answer to justify it's incorrect answer is another thing. A decent example between knowledge and understanding. It's not that a child would get the answer right, it's that the child would easily identify it's self inconsistency were it to give the same explanation.
As for gpt3, we have to remember that this is quite a different model than any of the chat versions (typically including gpt4). It was never trained through RLHF, which introducers many more biases as it dramatically changes the latent distributions. GPT base is often jokingly called a babbler, as these are more word prediction models. The chat aspect changes things. But I wouldn't expect anyone not deep in the literature to understand why these are extremely important differences just the same way I wouldn't expect an average person to understand why there's a big difference between a flat head screw and a slotted screw drive.
I don't want to shut down a conversation through authority (it doesn't exist on HN) but I have to state that it's difficult to go down this path without bringing up a lot more background material. We have to really dig into theory of mind, cognition, as well as get nuanced about NLP and LLMs in general. That's far too cumbersome than I'm willing to write in comments. But these things are essential priors to make the arguments we are discussing here. And I truly mean that this is not something that can be learned from YouTube and quite difficult to learn on the Internet. These are difficult subjects to learn in even the best settings with lots of nuances that are brushed away in introductory materials but critical when we discuss the meat and potatoes.
I used two words that probably don't exist anywhere in the training data and asked it (ChatGPT-3) to compare the weight of a kilogram of one vs an lb of the other. It initially gets it wrong, but when asked to show its work, it's suddenly able to decide that a kilogram of something is heavier than a pound of something. Using understanding or reasoning, or stochastic parroting; whatever you want to call it. I'm not claiming it's AGI, or sentient, but it's also clearly not just a Markov chain.