Reasoning skills of large language models are often overestimated
news.mit.edu
news.mit.edu
When it was previously discussed here, it received a torrent of low-quality comments that can only be described as confirmation bias: commenters tried the examples verbatim (instead of introducing random perturbations to get around model updates) in order to disprove the thesis.
It's sad that there are "engineers" out there so blinded by their own (wishful thinking, vested interests) that can not accept the obvious.
[1] https://arxiv.org/abs/2308.03762
[2] https://medium.com/@konstantine_45825/gpt-4-cant-reason-adde...
As it stands every publication on this subject is filled with a lot of people that feel they need a polarised stance or they'll immediately be shot: Either they say LLMs are fundamentally broken and can't do shit (this is ridiculous) or they say that LLMs are self-aware and we should give them every single task ever (also ridiculous). It is really tedious to read the same dumb shit every single time.
Not necessarily AGI or average human level reasoning, but just any form.
a man with a wolf and a cabbage wants to cross a river/ There is a boat but the man can only bring one more item beside himself at a time. What are the steps the man needs to do in order to cross the river?
The response was:To solve the puzzle where a man needs to cross a river with a wolf and a cabbage, while ensuring none of them are left together unsupervised (since the wolf will eat the sheep, and the sheep will eat the cabbage), follow these steps:
1.Man takes the sheep across the river (leaving the wolf
and the cabbage on the original side).
2.Man returns alone.
3.Man takes the cabbage across the river.
4.Man leaves the cabbage on the other side and takes the sheep back
5.Man takes the wolf across the river.
6.Man returns alone.
7.Man takes the sheep across the river.
8.This way, at no point are the wolf and the sheep left alone together, nor are the sheep and the cabbage left alone together.
at which I prompted: I didn't mention a sheepThe response was corrected but wrongly:
Oops, my mistake! Let’s solve the puzzle correctly with a man, a wolf, and a cabbage:
1.Man takes the wolf across the river (leaving the cabbage on the original side).
2.Man returns alone.
3.Man takes the cabbage across the river.
4.Man leaves the cabbage on the other side and takes the wolf back.
5.Man takes the man across the river.
6.Man returns alone.
7.Man takes the wolf across the river.
8.Now, both the man and the wolf are across the river, with no danger to the cabbage.
At step 3 the problem was solved but it kept on goingParlor tricks are lurking everywhere, and especially in ChatGPT, because of it's high profile and high stakes.
The only way for LLMs to get this right is to be trained on all possible paths of all possible problems, which is not very feasible. That's not really reasoning, you may as well calculate all permutations and add it to a DB and call it a day.
These models store these common patterns and then repeat slight variations of them as that is natural flow of the tokens. Humans do it similarly with the same problem if they have heard it before.
Challenge is that coming up with actually novel reasoning problems is not that easy. So generating one to test models is non-trivial task.
It's like debating your political other. You can often predict what they will say not because you understand their reasoning, but simply because you've seen multiple times how they react to a given argument in the past.
I mean, we'd get a similar outcome if we first give a human child this problem, and then subsequently hand them an ice-cream. "Clearly the child isn't reasoning or creating the way that adults reason, it's off eating an ice cream". :-P
We can debate whether the machine is smart, or whether natural language requires some level of reasoning as part of the <syntax>. (compare: the type system in rust) . We can even debate whether knowing natural language thus makes the machine smart. %-/
--
edit: Turns out GPT-4 (plain) can solve the question if you warn it upfront, and Claude actually recognizes the trap and then solves it. So I guess we've demonstrated that both AI and humans can jump to conclusions based on incomplete or assumed information.
Geeeeh, 4o really took a lot of work. It's like getting it to win at tic-tac-toe. It CAN be done, and possibly even repeatably, but it's non-trivial.
Try removing the wolf such that there is only a man and see how well it does. I did with but without 4o and it kept on making unnecessary trips where all could be solved in one step.
Try removing the wolf but add the sheep and the cabbage.
It seems to me that if falls of the rails rather quickly.
I tell you, 4o is like a child and ice-cream. You need to eliminate any hint of the ice cream before it'll even deign to do what you ask. %-/
--
Ok fine. I ran out of 4 and 4o responses, so instead I actually got GPT-3 to solve the man and the cabbage problem. But it was tough again, since 3 is bad at meta-evaluation for sure.
I keep getting the feeling that I'm doing "maths" with natural language instead of mathematical notation or code.
Not helped by the tendency of the AI to want to hare off into the wide blue yonder instead of stop and consider its own answers.
I shan't bother you with the transcript unless you really insist. (I had to convince it to check its own answers for logical deficiencies. I can't quite rule out a kluger hans effect[1] , though that's one heck of a kluger hans if so, having written out the entire sequence :-P )
--
Ok so a friend of mine tried this question with Claude 3.5 Sonnet, and Claude was almost offended by the question, and solved it perfectly. [2]
I don't have Claude, but I do realize that the web version of GPT-4 is actually a little smarter than the web version of 4o. Now that I have some 4 credits again:
https://chatgpt.com/share/e60d5237-c59d-442c-a59c-28a0dc12bd...
It got the answer wrong at first because it was over-eager, but when I told it it was wrong, it then actually took the correct approach.
--
Score so far:
GPT-3.5 : Can't manage the cabbage problem without a lot of help
GPT-4 : Needs to be told that it's walking into a trap, but then can solve it pretty much on its own after that.
GPT-4o: Fails. (I didn't try walking it through the reasoning, 4o is too tiring for me)
Claude-3.5: Insulted by the simplicity of the question, answers it, points out that this is not the man-wolf-goat question, and asks if it should answer that one too?
--
Mewpmewp2 figured out a way to avoid pattern matching, this solution actually works on all the GPTs I tried it on.
* https://news.ycombinator.com/item?id=40947296
* https://chatgpt.com/share/c3f744e4-199a-4c45-a332-c40d505326...
==
[1] Here the part of the effect is : "Keep asking, and stopping because I got the right answer"
[2] https://usercontent.irccloud-cdn.com/file/01Al3FZQ/172080080... <- extra edit: Claude 3.5 answer
https://chatgpt.com/share/2ac767d9-2a6b-4631-b071-7266e001a1...
My conclusion with LLMs is "Close but no cigar" when it comes to reasoning deviated from the script. Maybe it will get better but something has to change to make this work, more parameters or training will only make it appear to work better but still fail.
Even flawed I still find LLMS and GenAI useful for some tasks such as brainstorming though.
https://chatgpt.com/share/c3f744e4-199a-4c45-a332-c40d505326...
This one I used objects A, B and C, where A and C must move together with B.
https://chatgpt.com/share/c3f744e4-199a-4c45-a332-c40d505326...
Do you think that is also from a trained dataset?
Because I wonder if it might have some sort of internal router in mind where certain signals will trigger it to pattern match to a very common problem whereas if it's not overfit there it will try to fallback to some form of "reasoning".
* GPT-4o can solve the problem with this design.
* GPT-3 in my case actually discovered a lacking constraint, and transports A and C all at once. (I needed to add that B can only carry 1 object at a time)
https://chatgpt.com/share/9aa2efa1-ac89-4412-86b3-47fbce4107...
Myself, I think that's a terrible false analogy, and also have a higher bar for anything to be called artificial "intelligence."
These programs shouldn't be using personal pronouns or statements that imply cognition or esp reasoning. It's deceptive.
So I concur that the apologists are wrong about what the human brain is doing, and therefore wrong about the similarity to LLMs.
Also, generally, while humans can be really bad at reasoning, when they're shown why they're wrong they're able to retrace their steps and fix -- assuming they're not emotionally attached to some outcome. I'm terrible at math, but if I slow down and have it explained, I'll get it (and then forget it.) LLMs... they don't work that way. Tell them how they're wrong or to explain how they got to some result... forget it, they can't retrace and fix.
GPT-4o was a lot of work. GPT-3 is fast but that was a long conversation catching all its errors (so not posted, but I can if you insist). 4 caught its error almost immediately. Claude-3.5 was on to us from the get-go and didn't need correcting.
But I think we are only scratching the surface of achieving this. Based on my admittedly casual observation of current model capabilities you need a model that's fairly big (closer to 70B than 7B) and trained far beyond what's implied by the Chinchilla metric.
Previous discussion: https://news.ycombinator.com/item?id=40899309