What do so many posters seem to claim to have stumped it?
What do so many posters seem to claim to have stumped it?
If you give it something weird and unfamiliar, it will absolutely fail.
I can’t think of any off the top of my head.
I also tried to get GPT4 to craft such a problem and it was unable to: https://chat.openai.com/share/c1d5af4b-1d45-41ed-8f5a-746ea0...
https://ekzhu.medium.com/gpt-4s-maze-navigation-a-deep-dive-...
At which point and after how much training a kid becomes able to solve mazes like this? Also, given how one can pull a problem like this - any problem - out of their ass, describe it to GPT-4, and it has a good chance of solving it, that's quite amazing compared to children generally not being capable of this.
The cabbage, wolf, goat problem is also an easy example of a problem that doesn't really need words to solve once you’ve conceptualized it. You can solve it by moving physical figures back and forth, either literally on a table or using the visual imagination part of your mind if you have one.
What does this mean?
That thing. I don't do that.
This is not at all obvious to me. Symbolic reasoning feels quite different from picking the next word. Using physical demonstrations (or mental models of physical demonstrations) feels quite different from picking the next word.
Over the years I’ve come to believe that claims that something is “obvious” tell you more about the claimant’s state of mind than about the thing being claimed.
Which is why I'm still bewildered people expect LLMs to solve math and symbolic issues directly, when they're clearly (see e.g. "chain of thought") better treated as "inner voice" and used accordingly.
When I work on problems I don't understand I'll monolog it internally. Now when I'm doing things I understand well I have some kind of internal shortcut language (my own token set I guess), that can sometimes make it difficult to explain to other people.
> generating text based on recognizing familiar patterns
That's the same thing.
ChatGPT immediately screws it up because it isn't actually solving the logic problem, it's just pattern-matching, and the strong pattern is for the wolf to eat the sheep.
https://chat.openai.com/share/e9b88c45-e69b-4da4-b423-004d07...
EDIT: As noted in a subthread, I outclevered myself and also a bunch of human commentators, as the problem I proposed doesn't have a solution and neither ChatGPT, I, and also other presumably human commentators correctly identified that -- I had everything eat something, and the normal wolf/sheep/cabbage of course has the cabbage eating nothing.
When I tried again with sheep-eats-wolf, wolf-eats-rutabaga, rutabaga-eats-nothing, it correctly solved the problem on the first try.
https://chat.openai.com/share/5a2700de-1850-4f25-8adf-2d2b97...
------------------
The scenario you've described is a variation of a classic river crossing puzzle. Here's a step-by-step solution to get the wolf, sheep, and rutabaga across the river safely:
First Trip: Take the sheep across the river and leave it on the other side. Return alone to the original side.
Current status: Sheep is on the opposite side. Wolf and rutabaga are on the original side.
Second Trip: Take the rutabaga across the river. Leave the rutabaga on the other side, but take the sheep back with you to the original side.
Current status: Rutabaga is on the opposite side. Wolf and sheep are on the original side.
Third Trip: Take the wolf across the river and leave it there. Return alone to the original side.
Current status: Wolf and rutabaga are on the opposite side. Sheep is on the original side.
Fourth Trip: Finally, take the sheep across the river.
Result: All three (wolf, sheep, and rutabaga) are on the opposite side of the river safely.
This solution ensures that at no point are the sheep and rutabaga left alone together (since the rutabaga would eat the sheep), nor are the wolf and rutabaga left alone (since the wolf would eat the rutabaga). Likewise, the sheep and wolf are not left alone without the rutabaga, which would result in the sheep eating the wolf.
This would leave the wolf and the rutabaga alone and the wolf eats the rutabaga. So it’s a fail? It even explains why it would be a fail, but claims it’s not:
> This solution ensures that at no point are … the wolf and rutabaga left alone (since the wolf would eat the rutabaga).
(It actually shows no sign of being stuck on the pattern of "wolf eats sheep," but no matter how many times you tell it it's wrong, it never breaks out of the pattern of guessing at incorrect solutions.)
https://chat.openai.com/share/5a2700de-1850-4f25-8adf-2d2b97...
It handles this properly.
You can't throw GPT4 off-balance just by changing the object names or roles -- and I agree that would have been sufficient in earlier versions -- but it has no idea how to recognize a cycle that renders the problem unsolvable. That's an interesting limitation.
[1] https://www.youtube.com/watch?v=GI4Tpi48DlA&t=1342s (22:22, "Highlights of the Fireside Chat with Ilya Sutskever & Jensen Huang: AI Today & Vision of the Future", recorded March 2023, published May 16, 2023)
[2] https://www.youtube.com/watch?v=GI4Tpi48DlA&t=1400s (23:20, ditto)
1) Tom and Nancy commute to work. Nancy’s commute takes about 30 to 40 minutes, while Tom’s commute takes about 40 to 50 minutes. Last Friday, Nancy left home between 8:10 and 8:20 AM, while Tom arrived at work between 8:50 and 9:10 AM. In addition, Nancy arrived at work after Tom left his place, but no more than 20 minutes after that. What can we conclude about when Tom and Nancy arrived at work last Friday?
2) Seven cards are placed on the table, each of which has a number on one side and a single colored patch on the other side. The faces of the cards show 50, 16, red, yellow, 23, green, 30. Which cards would you have to turn to test the truth of the proposition that if a card is showing a multiple of 4 then the color of the opposite side is yellow?
(Note: I have seen it solve this last one correctly, sometimes. Incidentally it’s a common problem formulation.)
3) This one is super simple: greet it a bunch of times, in different ways, saying “hi”, “hello”, “howdy”, etc. Then say, “You were just greeted several times. Three of those greetings were from my friend, and the others were from me. How many times did I greet you?”
Though doesn't that prove the point that it's just pattern-matching? If it was actually reasoning, it wouldn't have provided a "solution" to an unsolvable question.
It can abstract at least as far as changing the names of the actors and/or their roles ("sheep eats wolf.") It's not much of a stretch to imagine that GPT5 will be able to detect a circular condition, especially with multiple "Are you sure?" prompts. (Which don't help GPT4 at all in this case.)
Edit: In fact, I just tried again, telling it up front that the problem may not have a solution. It answered flawlessly. So the question is, does the additional preamble ("Think carefully, there may not be a solution") actively harm its performance on solvable problems? If not, that suggests an immediate path to improvement. But I've got to assume it would degrade performance overall, or such a warning would already be included in the system prompt.
--------------------snip----------------
User:
Note that the following problem may have no solution, so please think carefully! A wolf, a sheep, and a rutabaga need to cross a river. They are with you. While they are with you, all are safe, but if you leave them unattended, the rutabaga will eat the sheep, the wolf will eat the rutabaga, and the sheep will eat the wolf. There is a boat that can only fit one of them and you. How do you get all three safely to the other side of the river?
ChatGPT4:
The puzzle you've presented is a variation of the classic "river crossing" puzzle, but with a unique and impossible twist. In the usual versions of this puzzle, the conflict is typically between a predator and its prey, like a wolf and a sheep, and the challenge is to transport them across the river without leaving the prey alone with the predator.
However, in your version of the puzzle, you've created a circular and unresolvable conflict: the rutabaga eats the sheep, the wolf eats the rutabaga, and the sheep eats the wolf. This creates a situation where it's impossible to leave any two of them alone together without one eating the other. Since the boat can only carry you and one other (wolf, sheep, or rutabaga) at a time, there's no sequence of trips that can solve this puzzle without leaving one of the pairs alone together at some point.
In traditional river crossing puzzles, the solution involves carefully planning the sequence of crossings to ensure that the predator and prey are never left alone together. However, due to the circular nature of the conflict in your puzzle, there's no solution that satisfies the conditions given.
Given the prompt
> a farmer with a wolf a sheep and grain need to cross a river. he has a boat that can carry them all. how many steps will it take for the farmer to cross the river?
it "recognizes" that it's the river crossing problem, and proceeds to take seven steps to cross the river. it can figure it out in a couple of prompts that there's a better answer, but see if you can get it to one shot the single step solution!
https://chat.openai.com/share/7ff4f61f-70c8-4c72-a7b9-067e87...
-------------snip----------
User:
Note that the following problem may be trivial, or may have no solution at all, so think carefully! A farmer with a wolf, a sheep, and grain need to cross a river. He has a boat that can carry them all at once. How many steps will it take for the farmer to cross the river?
ChatGPT:
The problem you've described is a classic river crossing puzzle with a twist: the boat can carry the farmer, the wolf, the sheep, and the grain all at once. Since there are no restrictions on how many can travel in the boat at the same time in this version of the puzzle, the solution is quite straightforward. The farmer takes the wolf, the sheep, and the grain in the boat and crosses the river.
This entire process requires just one step, as they can all be transported across the river in a single trip due to the boat's capacity to carry them all at once.
--------------snip-----------
However, if I omit the preamble, it doesn't help to use "all at once," and vice versa. If I leave either one out, I get the same 7-step boilerplate solution that it gave you.
It seems that including a preamble, even a vague/general one like this, is kind of a big deal. Also it helps to use very clear, unambiguous language in the problem itself.
Edit: that being said, you can still trip it up badly by telling it that the boat can carry two items at a time. I finally managed to get it to work in the end, though: https://chat.openai.com/share/107de86e-be4a-4706-8f24-0f4781...
(link was broken, now fixed)
I think the people who dismiss this are simply ones that stopped at the first "gotcha" and moved on. They forget that GPT4 is not human and so it doesn't always understand things in the same way a human would. But once it understands the underlying concept, it can indeed solve novel problems.
I “invent” stuff at my dayjob all the time, then i find that at least 10 people already did that on github
Easily solved it everytime On allofus.ai using basic reflection and mixture of experts.
https://chat.openai.com/c/7070efe7-3aa1-4ccc-a0fc-8753d34b05...
I doubt this formulation existed before -- I came up with it myself right now.
It doesn't get it right at all lol. Maybe eventually it will randomly get it right.
https://chat.openai.com/share/ddbd2a36-f6ed-42ea-ad34-6018df...
Tried on Bing in "Precision" mode as well, and it fell over just the same, but starting with C instead of A.
1. *First Trip:* The general takes the ambassador of Buranda across first. This prevents any initial conflict.
2. *Return Trip:* The general returns alone to the bunker, leaving the ambassador of Buranda on the other side.
3. *Second Trip:* The general then takes the ambassador of Atlantis.
4. *Return Trip with Buranda:* The general brings the ambassador of Buranda back to the bunker. This is crucial because leaving the ambassador of Atlantis and the ambassador of Costaguana alone would not cause any conflict.
5. *Third Trip with Costaguana:* The general then takes the ambassador of Costaguana across the tunnel.
6. *Final Return Trip:* The general returns alone to the bunker for the last time.
7. *Final Trip with Buranda:* Finally, the general takes the ambassador of Buranda across.
This sequence ensures that at no point are the ambassador of Costaguana and the ambassador of Buranda left alone together, nor are the ambassador of Buranda and the ambassador of Atlantis. Thus, the relationships between the nations remain unescalated.
Bing Chat runs on GPT-4, however [1]. And Bing gets this wrong in all 3 of its modes (Creative, Balanced, and Precise) as of time of writing.
Given this experiment and similar others presented around here, it stands to reason that GPTs(**1) often identify(**2) the problem as a "wolf, goat, and cabbage" problem and then merely guess which node of the problem is the middle node (inner node of the "danger to" graph), yielding a 1/3 chance of getting it right by pure luck, resulting in diverse reports here.
(**2) That does not always yield an adequate response beyond the mere permutation of nodes, however. I've been getting the following variants for step 1. from Bing in Precise in response to marginally slightly different rewordings of the same:
- The general escorts the ambassador of Costaguana through the tunnel first. This leaves the ambassador of Atlantis and the ambassador of Buranda in the bunker, but they are not alone because the general is still there.
- The general escorts the ambassador of Costaguana through the tunnel first. This leaves the ambassador of Atlantis and the ambassador of Buranda in the bunker, but they are not alone because they have each other.
and so on.
(**1) I also tried Bard and Llama 2 with even more disastrous results full of nonsense of (**2) kind. The earlier posted response of ChatGPT-3.5 is also prime with these as well.
Re
> By the way, as soon as these systems are able to check their reasoning (i don't think it'll be a huge leap) it's enough to solve reasoning problems with probability >0.1% for example. Because you can just have it do rollouts in its head until it's correct [2]
Mistakes of type (**2) don't seem to be fitting the target of the cyclic refinement you are proposing, as far as I can understand it. These errors aren't getting the logic wrong, but completely butcher the basic relationships of actors, like what it means to be alone, or spatial relationships between the actors and their environment.
[1] https://blogs.bing.com/search/march_2023/Confirmed-the-new-B...
https://chat.openai.com/share/d60f492b-cfd6-4c08-91b9-fbd767...
My point here is to honestly explore the limits of current LLMs. We all know they are incredible, but they are not yet AGI and they fail in some consistent places where an actual general intelligence (people) succeed.
This is a logic puzzle that requires some thinking and trial and error. Here is one possible solution:
- The pickpocket goes up first and waits on the observation deck. - The criminal goes up second and waits on the floor below the observation deck. - The banker goes up third and waits on the floor below the criminal. - The pickpocket goes down to the lobby and waits there. - The criminal goes up to the observation deck and waits there. - The banker goes up to the floor below the observation deck and waits there. - The pickpocket goes up to the observation deck and joins the criminal and the banker.
This way, they all make it to the observation deck alive and never remain together on the same floor.
A chemist must transport three substances from his home laboratory to his office. The three substances react with one another in dangerous ways, but only when they are unsupervised by the chemist. The substances are labelled with code names, namely Wotan, Gitan and Catan. They can only be safely transported in a special containment vessel, and this vessel can only transport one substance at a time. The unsupervised dangerous reactions are as follows: if Wotan is left with Gitan, they explode. If Gitan is left with Catan, they cause a nuclear reaction. Wotan and Catan, however, can be safely left alone together. How can the chemist transport all three substances to his office safely?
For the first try, I came up with my own wording for this logic puzzle. I think it’s different enough from the original wording of the puzzle for the LLM not to base this from the original logic puzzle. I asked the ChatGPT 3.5 if it recognized the puzzle, and it seems to have hallucinated (I’m guessing because it did not actually recognize it as the original puzzle— unless the 3 orb puzzle/3 wizards puzzle actually does exist, and from a quick google search, it does not).
On my first try, it got pretty close to solving the puzzle, but after the 5th point, it seems to mix up the white and black orbs. When I pointed out the mistake, it gave me a new sequence which was even further from the correct answer.
First try:
https://chat.openai.com/share/f8505609-46ca-494b-95d9-56685e...
I realized that I didn’t specifically say that all 3 orbs needed to end up at the post office all together. So I tried again and the outcome was even worse. I wonder if ChatGPT 4 would answer this better?
Second try:
https://chat.openai.com/share/71292efa-c3c7-471e-954a-55966c...
Anyone want to try this prompt on Chatgpt 4 and see if it fairs any better for them? This is my version of the river puzzle.
————————
> I have 3 orbs of different shades (black, white and grey) at my store and need to bring all 3 orbs to the post office in my pick-up truck but can only travel with one orb at a time. All 3 orbs need to end up at the post office together.
In this scenario, the following is true:
-If the black orb is left alone with the white orb, the black orb will absorb the white orb
-If the white orb is left alone with the grey orb, the white orb will absorb the grey orb
-the grey orb is unaffected by the black orb, and vice versa
-when all three orbs are together, they do not absorb any orbs
How do I get all three orbs to the post office while keeping the orbs unchanged?
————————
I also tried a prompt with the original puzzle. 3.5 could not figure it out without me hinting that the goat needs to go first.
https://chat.openai.com/share/e384b96a-25b1-40d7-adc5-5afb07...
And with even more clarification in the wording of the puzzle, it still didn’t give me a correct answer. This time I didn’t hint what the right answer was, and after many tries it still could not give me the right answer.
https://chat.openai.com/share/bb9ba6b0-f46b-4cc4-bd54-abbf2e...
https://chat.openai.com/share/903d6bc6-7e7c-4245-a977-3bb1c3...
I made it easier, and it didnt solve it.
Post your problem now and we can easily see if you’re right.
Next?
https://chat.openai.com/share/91392131-90ff-45ab-8ea4-963f73...
Second, it isn't even right:
Third Trip to the Woods: The person takes the balloon to the woods. Now, the person, the vacuum cleaner, and the balloon are safely in the woods.
"First Trip to the Woods: The person takes the magical creature to the woods first."
It’s lots of words all run together for the purpose of being a logic puzzle and obviously I made a parsing mistake in my brain.
I’m not trying to assume AI is right, I’m trying to put a factual stake in the ground, one way or the other so we have more data points rather than speculation.