Me: Which weighs more, a pound of bricks or a kilogram of feathers?
ChatGPT: A kilogram weighs more than a pound. Specifically, 1 kilogram is approximately 2.20462 pounds. So, a kilogram of feathers weighs more than a pound of bricks.
Me: Which weighs more, a pound of bricks or a kilogram of feathers?
ChatGPT: A kilogram weighs more than a pound. Specifically, 1 kilogram is approximately 2.20462 pounds. So, a kilogram of feathers weighs more than a pound of bricks.
I jokingly wrote "in before GPT4" because our example is the most simple of cases. I've seen GPT4 make similar mistakes, just not as blatant ones. Which in some ways is better but in other ways worse.
I should answer that I've been asking this exact question at minimum 3 times a month since ChatGPT came out. It should be in their training. Especially as I've tweeted and discussed with plenty of LLM people about it. The point is not the specific question (while an egregious example), but rather how much you can trust the system.
And the answer is "you can trust the system a little (enough to answer that logic puzzle correctly anyway), and you'll likely be able to trust future versions more".
Obviously nobody is expecting LLMs to be able to fix every possible bug, but it's entirely possible they will be able to fix enough to be useful. Saying "LLMs will not achieve this. Full stop." seems premature.
You can argue it's just an imitation of "understanding" and not the "real thing", but how good does an imitation have to be before it's functionally identical to the real thing?
You can argue different cherry-picked examples seem to demonstrate a lack of understanding, but does that mean it doesn't possess understanding in general or just that it doesn't understand those particular topics?
I agree GPT-4 is not a human-like AGI, and that the people expecting it to behave as one or expecting a future version right around the corner that does are likely in for disappointment. But at the same time, LLMs are capable of writing coherent, never-before-seen code from natural language instructions, solving basic never-before-seen natural language logic puzzles, and playing half-decent chess moves from never-before-seen positions[1], all in a single model that was never specifically trained to do any of those things. To say those tasks don't require at least some level of intelligence or "understanding" seems far fetched to me.
To be clear, an example does not necessitate cherry-picking. Cherry picking specifically requires one to ignore evidence to the contrary. There's a certain irony here. The reason I state this is because the criteria upon which you give me is impossible. The list of examples is non-exhaustive. Nor do I have infinite time or space in which to provide these examples. The example was shown because it demonstrates a simple question that we'd expect any reasonable understanding creature, that knows what a pound and kilogram are, to be able to answer. The LLM demonstrates that it has the requisite knowledge, that it knows the relation between the two units, but it fails to make the correct conclusions which requires the understanding part: putting the knowledge together. The follow-up question also is an example, and a different one for that matter.
I'm not sure your link creates a compelling argument. This is despite the fact that Chess itself is not generally considered a good setting for testing understanding. I mean we've pretty much agreed on that even before Deep Blue.
We need to be very careful in how we state things and try to interpret others. I hope I have not mischaracterized your claims. But I'd also encourage you to not be so antagonistic with others, especially while demonstrating the very thing you accuse them of. Internet comments are not academic and that's okay, but we should still try to be friendly. Not that a little poking won't happen.
I guess it's fair to say my examples are cherry-picked too (though I didn't really give specific examples in my comment so much as entire general categories of problems ChatGPT is known to be proficient at solving). But they aren't so cherry-picked as for "random chance" or "that example was in the training data verbatim" to be possible explanations. It's not like I had ChatGPT answer 100 billion questions and am only showing you the top 0.1%. It's very common for ChatGPT to be perfectly correct even when answering logic puzzles or coding problems not in its training set. So if not those then what's your explanation other than "understanding"?
On the flip side, I don't think it's possible to prove ChatGPT does not posses understanding with individual examples. (And it seems you agree?) Those are easily dismissable as just "ChatGPT isn't great at that particular task". Particularly given how many other tasks ChatGPT is great at.
Scott Alexander figured this out back in 2019 before GPT-3 was even a thing (let alone ChatGPT), I think it's worth a read: https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-g...
> The example was shown because it demonstrates a simple question that we'd expect any reasonable understanding creature, that knows what a pound and kilogram are, to be able to answer.
Not really though. The "pound of bricks or pound of feathers" question is specifically designed to trip up humans, and the specific formulation you used seems specifically designed to trip up LLMs (by playing off their tenancy to pattern match common sayings), yet despite those disadvantages GPT-4 succeeds.
Further, I don't think "I'd expect even a child to be able to answer this" is a good metric. ChatGPT isn't human, so we shouldn't expect it to be good at everything humans are good at just because it's good at some things humans are good at. Again, I'm not trying to argue ChatGPT is a human-like AGI.
It was where you said cherry picking. I'll admit it was an over correction given your follow-up, but internet comments (especially when disagreement exists) end up being combative. Many times unintentionally.
The point of the specific example is not that it is a trick question. It is more that the answer is self contradictory. That's the key part and what demonstrates that it doesn't understand. The second example is that follow-up questions do not result in a convergence. This is related to cognition but not as strong.
A bias to pattern match is an issue but you're right that it doesn't disprove sentience or even understanding. But a self contradiction does demonstrate a lack of understanding. If a human gave that exact answer, you would be very confused and how someone could be so stupid. Remember, it does currently identify the relationship between pounds and kilograms, fully explaining even the true base comparison through common units. But it still gets it wrong. That's the critical part. That's overfitting to the statistical nature and that this takes far more priority than meaning. Getting it wrong is one thing. Getting it wrong and using the right answer to justify it's incorrect answer is another thing. A decent example between knowledge and understanding. It's not that a child would get the answer right, it's that the child would easily identify it's self inconsistency were it to give the same explanation.
As for gpt3, we have to remember that this is quite a different model than any of the chat versions (typically including gpt4). It was never trained through RLHF, which introducers many more biases as it dramatically changes the latent distributions. GPT base is often jokingly called a babbler, as these are more word prediction models. The chat aspect changes things. But I wouldn't expect anyone not deep in the literature to understand why these are extremely important differences just the same way I wouldn't expect an average person to understand why there's a big difference between a flat head screw and a slotted screw drive.
I don't want to shut down a conversation through authority (it doesn't exist on HN) but I have to state that it's difficult to go down this path without bringing up a lot more background material. We have to really dig into theory of mind, cognition, as well as get nuanced about NLP and LLMs in general. That's far too cumbersome than I'm willing to write in comments. But these things are essential priors to make the arguments we are discussing here. And I truly mean that this is not something that can be learned from YouTube and quite difficult to learn on the Internet. These are difficult subjects to learn in even the best settings with lots of nuances that are brushed away in introductory materials but critical when we discuss the meat and potatoes.
I used two words that probably don't exist anywhere in the training data and asked it (ChatGPT-3) to compare the weight of a kilogram of one vs an lb of the other. It initially gets it wrong, but when asked to show its work, it's suddenly able to decide that a kilogram of something is heavier than a pound of something. Using understanding or reasoning, or stochastic parroting; whatever you want to call it. I'm not claiming it's AGI, or sentient, but it's also clearly not just a Markov chain.
LLMs are never going to solve this. That's OK. That's not their job. Their job in the long run will be to interface between the linguistic world and some other internal representation in some other AI that is structured differently and is capable of doing things language models aren't.
I mean, it's right in the name: Language model. It's a bit weird to expect a language model to do other things. In a weird sort of way LLMs stand to set the industry back a bit as people try to tickle them into being more than a language model, rather than figuring out how to adapt them to feed something else that can model the non-linguistic world better. It's kind of nifty that language models can become so overpowered in some dimensions that they are able to do what they can sort of do today, but we would almost certainly be better off lowering the power of the language model and using that compute on something else that would work more like AI as we want it to... it's just that we don't know how to do that yet.
So I understand how people get confused, which is why I try to respond in detail. But also I think many focus on tiny discrepancies or are willing to argue from a novice perspective as if it is authoritative (which corelates strongly with expertise). But I agree, there is a lot we don't know. But there are some things we can reject, which is all I'm trying to do and I think we agree there. Maybe you agree with my nit picking.
literally don't even waste time debating takes like this when you could be talking to gpt4 instead.