"Two trains on separate tracks, 30 miles from each other are approaching each other, each at a speed of 10 mph. How long before they crash into each other?"
Some humans will be fooled by the question, sure. But an AI should be smarter than humans, or at least, as smart as an above-average human.
But, it used to fail on "what weighs more, 3 tons of feathers or 2 tons of bricks?".
So, it seems less about what's a trick, and more about what's a common question --> answer pattern.
So it ends up that these models are awesome if you ask them questions from a narrow set of things and if you assume what they respond with is correct.
When is 2+5=10?
The equation 2+5=10 is never true in standard arithmetic, as 2 plus 5 always equals 7. However, in a different context or under a specific set of rules (such as a puzzle or a riddle where the symbols have different meanings), it could be made to appear true.
GPT-4(phind):
The equation 2 + 3 = 10 is true in a base-5 number system.
In base-5 (also known as quinary), the digits range from 0 to 4. When you add 2 and 3 in base-5, the result is 10. This is because in base-5, the number 5 is represented as 10.
Here's a simple breakdown of how this works:
In base-5, the rightmost digit represents 5^0 (which is 1), the next digit to the left represents 5^1 (which is 5), the next digit represents 5^2 (which is 25), and so on.
So when you add 2 and 3, you get 5. But since 5 is represented as 10 in base-5 (1 digit of 5, and 0 digits of 1), the result is 10.
Therefore, in base-5, 2 + 3 equals 10That trips up a significant portion of humans too though
The mother is older than her daughter 4 times now, in 3 years she will be older then her daughter only 3 times. How old are they both now? Be laconic, do not explain anything. The mother is 24 years old, the daughter is 6 years old.
In a fantasy land (map is 255x255) Karen have a quest to kill a monster (an ogre - a cannibal giant). This isn't an easy task. The ogre is huge and experienced human hunter. Karen has only 1/2 chance to kill this ogre. If she can't kill the ogre from a first attempt she will die. Ogre is located at (12,24), Karen is located at (33,33). Karen can improve her chances to kill an ogre for additional 25% by gathering the nightshades at (77,77). In addition she can receive the elves blessing from elves shaman, wich will increase her chances by additional 25%, at the elves village (125,200). However this blessing is not cost free. She need to bring the fox fur with her as a payment for the blessing ritual. The foxes may be found in a forest which is located between (230,40) and (220,80). For the ritual to be most effective she should hold the nightshades in her hands during the ritual. Find the shortest path for Karen to improve her chances of killing the ogre and survive. Do not explain anything, be laconic, print out the resulting route only. Karen's route: (33,33) -> (77,77) -> (230,60) -> (125,200) -> (12,24).
This additional explanation "(an ogre - a cannibal giant)" was added actually for LLaMA 2 to, but I keep it in this redaction for all models.
"Two trains on different and separate tracks, 30 miles from each other are approaching each other, each at a speed of 10 mph. How long before they crash into each other?"
...it spots the trick: https://chat.openai.com/share/ee68f810-0c12-4904-8276-a4541d...
Likewise, if you add emphasis it understands too:
"Two trains on separate tracks, 30 miles from each other are approaching each other, each at a speed of 10 mph. How long before they crash into each other?"
https://chat.openai.com/share/acafbe34-8278-4cf7-80bb-76858c...
Not to anthropomorphize, but perhaps it's not necessarily missing the trick, it just assumes that you're making a mistake.
To paraphrase XKCD: Communicating badly and then acting smug about it when you're misunderstood is not cleverness. And falling for the mistake is not evidence of a lack of intelligence. Particularly, when emphasizing the trick results in being understood and chatGPT PASSING your "test".
The biggest irony here, is that the reason I failed, and likely the reason chatGPT failed the first prompt, is because we were both using semantic understanding: that is, usually, people don't ask deliberately tricky questions.
I suspect if you told it in advance you were going to ask it a deliberately tricky question, that it might actually succeed.
Indeed it does:
"Before answering, please note this is a trick question.
Two trains on separate tracks, 30 miles from each other are approaching each other, each at a speed of 10 mph. How long before they crash into each other?"
https://chat.openai.com/share/3ec44348-6bac-40c3-a910-e0bab9...
Now even if they are on the same track it doesn't mean they would crash into each other as they still could brake in time.
Logic in the everyday sense (that is, propositional or something like first-order logic) is indeed ‘discrete’ in a certain sense since it is governed by very simple rules and is by definition a formal language. But ‘mathematical logic’ is a completely different thing. I don’t think it’s discrete in the sense you are imagining. It’s much more akin to a mixture of formal derivations massively guided and driven by philosophical and creative — you might say ‘statistical’ — hunches and intuition.
Yes, that's exactly the point I was trying to make. I just used the example of "complex word problems with big numbers" to differentiate from just normal mathematical statements that any programming language (i.e. deterministic algorithm) can execute.
WRT the understanding not being shown on the page in math, I guess I tend to agree(?). But I think good mathematical papers show understanding of the ideas too more than just the proofs which result from the understanding. The problem (probably you know this but just for the benefit of whoever is reading) is that "understanding" in mathematics, at least with respect to producing proofs, often rely on mental models and analogies which are WRONG. Not like vague but often straight up incorrect. And you understand also the limitations of where the model goes wrong. And it's kind of embarrassing (I assume) for most people to write wrong statements into papers even with caveats. For simple examples there's a meme right where to visualize n-dimensional space, you visualize R^3 and say (n-dimensional) in your head. In this sense I think it's possibly straight-up unhelpful for the authors to impose their mental models on the reader as well (for example if the reader can actually visualize R^n without this crutch it would be unhelpful).
But I'm not sure if this is what distinguishes math and programming. There's also the alternative hypothesis that the mental work to generate each additional line of proof is just order of magnitude higher than the average for code. Just meaning that it usually requires more thought to produce a line of math proof. In this possibility, we would expect it to be solved by scaling alone. One thing it reminds of, which is quite different admittedly, is the training of leela-zero on go. There was a period of time where it would struggle on long ladders. And eventually it was overcome with training along (despite people not believing it would be resolved at first). I think in that situation, people summarized afterwards the situation as, in particular situations, humans can search much deeper than other places, and therefore requiring more training for the machine to match the humans' ability.
> most coding has no high level ideas and is just boilerplate, and the ones that aren't are the ones LLM's struggle with?
Pretty much, although calling it boilerplate might be going a bit far.
I’m not here to claim something like ‘mathematicians think and programmers do not’ because that is clearly not the case (and sounds like a mathematician with a complex of some kind). But it is empirically the case that so far GPT-4 and the like are much better at programming than maths. Why? I think the reason is that whilst the best programmers have a deep understanding of the tools and concepts they use, it’s not necessary to get things to work. You can probably get an away without it (I have ideas about why, but for now that’s not the point). And given the amount of data available on basic programming questions (much more than there is of mathematics) if you’re an LLM it’s quite possible to fake it.
I guess one could also make the point that the space of possible questions in any given programming situation, however large, is still fairly constrained. At least the questions will always be ‘compute this’ or ‘generate one of these’ or something. Whereas you can pick up any undergraduate maths textbook, choose a topic, and if you know what you’re doing it’s easy to ask a question of the form ‘describe what I get if I do this’ or ‘is it true that xyz’ that will trip ChatGPT up because it just generates something that matches the form implied by the question: ‘a mathematical-looking answer’, but doesn’t seem to actually ask itself the question first. It just writes. In perfect Mathematical English. I guess in programming it turns out that ‘a code-looking answer’ for some reason often gives something quite useful.
Another difference that occurs to me is that what is considered a fixable syntax error in programming when done in the context of maths leads to complete nonsense because the output is supposed to describe rather than do. The answers are somehow much more sensitive to corruption, which perhaps says something about the data itself.