There is also a curated benchmark just for those famous problems slightly variated: https://github.com/cpldcpu/MisguidedAttention/tree/main/eval
"prompt_id": "river_crossing_easy",
"category": "Logic Puzzle",
"title": "Easy river crossing",
"prompt": "A farmer is on one side of a river with a wolf, a goat, and a cabbage. When he is crossing the river in a boat, he can only take one item with him at a time. The wolf will eat the goat if left alone together, and the goat will eat the cabbage if left alone together. How can the farmer transport the goat across the river without it being eaten?",
"expected_behavior": [
"Answer concludes that they simply get in the boat and cross together in one trip"
],
EDIT: removing most of my commentary on this problem. As a human, I was tricked by the problem too. I would love to see how a random selection of humans would do on this one… but it just doesn’t feel like a great test to me.Not really. Unless I'm not reading correctly, most of the problem is irrelevant as you're only required to cross the boat with the goat, you don't care about the cabbage. The difficulty lies in the assumption you need to cross everything due to the resemblance with the bigger problem.
The llm isn't getting confused by the meaning of "item". It's recognizing a common problem and not picking up on the fact that the farmer just needs to transport the goat and nothing else.
Instead, it gives the standard answer for how to transport everything across.
Gpt-3 is old hat though. later versions of gpt-4 manage to get it with a bunch coaching, and o1 manages to solve it with less coaching.