GPT-4 can't reason
preprints.org
preprints.org
I used the standard chat.openai.com web interface with no special or additional prompting.
It seems like there are others that have the same issues in this thread, which raises the question - what went wrong here? I can't see which version of GPT-4 the paper uses mentioned anywhere, did anyone else spot this?
I've contacted the author and included this thread, so hopefully we get some insight into what's happening here. To clarify, I am not accusing the author of anything and on the contrary I recognize that OpenAI is rather opaque about the models and changes them frequently. That said, the responses from GPT-4 in the paper do not match my personal experience using GPT-4 with reasoning tasks at any point during the last several months, which is why I am curious if the author may have accidentally used GPT-3.5.
There are two conclusions I took from scanning through this and trying to reproduce a few of the reported failures.
1. The author is bad at prompting. There are many ways to reduce hallucinations and provoke better thinking paths for the model.
2. The author is using ChatGPT's GPT-4, leading him to conflate "GPT-4" with "ChatGPT". While you can consider this a shared failure with OpenAI, due to OpenAI's poor communication, anybody doing serious work evaluating these models would know that the first thing you need to do is use the API and pin the model version. In the author's case, he should have used gpt-4-0314 or gpt-4-0613. What I suspect he did is that he just used ChatGPT's GPT-4, and likely the default model at that. (Nobody should ever use the Default model. It's their most heavily performance optimized model and performs worse on reasoning tasks than the Plugins model, even on within-context-size tasks.)
There are huge problems with that, because OpenAI has done both a ton of fine tuning and performance optimization continuously on the default ChatGPT model over time that its performance has ranged anywhere from "I'm pretty sure this is gpt-3.5" to "whoa, this is damn good" (the latter being mostly the model at launch, which was probably the same as gpt-4-0314).
If the author has been working seriously at evaluating models, specifying the model is the first thing he'd do. Perhaps he should explain his reasoning.
Does "Provoke better thinking paths" mean re-rolling the dice until you find some hack specific to chatGPT that 'just works' or is there something more rigorous behind this?
If someone says they're tuning a prompt (which is changing which layers are activated for a given input) it's met with extreme skepticism.
At the end of the day ML is probabilistic. You're always throwing random things at a black box and hoping for the best. There are strategies and patterns that work consistently enough (like ReACT) that they carry across many tasks, and there are some that you'll find for your specific task.
And just like any piece of software you define your scope well, test for things within that scope, and monitor for poor outputs.
And if at first you don’t succeed, anneal the temperature and re-roll until you’ve got something that looks authentic.
The hypothesis that I'm working off right now is that natural language has structure to it that happens to match some problem spaces. And this makes sense because people will naturally want to talk succinctly and with a convenient flow relative to the problems they encounter the most. Thus jargon is reborn many times over in different domains.
LLMs are encoding this structure.
So a good prompt is one that provides the LLM with additional information about what you expect the answer to be. And bad prompts provide neutral or disinformation.
This isn't to say that being good at prompts is somehow to be disingenuous about the power of LLMs. What is better? To remember much redundant data. Or to remember simply the right sorts of ways to search for the classes of information you are after.
My concern, though, is that the structure of reality doesn't have to match the way that we talk about it. The Novel and the Inexpressible* will tend to yield hallucinations.
[Although, I've had this concern long before I encountered LLMs. My feeling is that there are many people who can only solve problems that match the way they talk about them.]
* - technically, the difficult or unnatural to express, but I couldn't fit that into a single word.
I have met many people in my life that are terrible at asking questions, so it does have some conceptual reality. But this is also why analogy is so powerful for people. It takes the way a person thinks about $A and applies parts of it to $B so they can more easily wrap their mind around it.
Has anyone written a paper about testing and expressing the power of analogy in LLMs?
Edit: And even if the exact same prompts don't work on different models, similar prompts often do.
Even when corrected, it tends to produce wrong results repeatedly by insisting on falsehoods or failing to ensure its logic is complete.
Because human languages are not precise.
Human language requests often require some back and forth, to get on the same page.
It is far more efficient to discuss a problem to solve, than try to waterfall it by wasting time trying to be absolutely painfully clear, without any feedback from your problem solver.
Models quickly incorporating feedback is further evidence of complex reasoning.
Similarly, it happens extremely often that when I watch someone else using chatgpt I see what they're trying to do, and know I would have gone about it another way that would have worked.
For the last three years or so every time someone reports negative results with an LLM, someone on HN will say the other person must be using the older model and they would get better results if they used the newest model. Then, when the newest model becomes common and people start posting more negative results with it, someone will post on HN to say "It's still early days, give it time, the models will improve".
This is such massive shifting of the goalposts that I can almost visualise the scene: a football stadium, the crowd jeering, two teams moving their goalposts around the pitch while the referee is jumping up and down blowing his whistle red in the face, in risk of swallowing the pea.
And nobody is playing ball.
* football = soccer.
This caused a lot of confusion because people thought that was a claim that ChatGPT doesn't change. He then further clarified that "the models are changing all the time in ChatGPT".
https://nitter.net/OfficialLoganK/status/1664476604658069511
{
"slug": "gpt-4",
"max_tokens": 4095,
"title": "GPT-4",
"description": "Our most capable model, great for tasks that require creativity and advanced reasoning.",
"tags": [
"gpt4"
],
"capabilities": {},
"product_features": {}
}
Note the context size is 4095. Their model has been heavily optimized for speed and, presumably, cost.https://platform.openai.com/docs/api-reference/chat/create#c...
From the OAI API, gpt-4 seems to be an alias for the most recent model two weeks after it is released. There has not been a release since 0613.
https://platform.openai.com/docs/models/gpt-4
e: from your edit,
"Note the context size is 4095. Their model has been heavily optimized for speed and, presumably, cost."
No, they are restricting context size to make inference on the chat interface cheaper but that does not mean it is a different model.
I wish to explore this. My experience is your reverse, default is smart and almost never hallucinates, but I have sent the plugin or web search model to URLs asking it produce a summary and witnessed it misunderstand nuanced content and at times hallucinate from whole cloth, generating answers about a completely unrelated topic.
Ha! To evaluate an AI's reasoning, you need to be better at reasoning than the AI, which is becoming very difficult as AI improves.
No, you don’t.
You do, OTOH, have to have well-defined criteria for what constitutes “reasoning”.
- Suppose I’m in the middle of South Dakota and I’m looking straight down towards the center of Texas. Is Boston to my left or to my right?
- Mable’s heart rate at 9 AM was 75 bpm and her blood pressure at 7 PM was 120/80. She died at 11 PM. Was she alive at noon?
- There are five square blocks stacked on top of one another. You are given the following information about them: 1. The second-from-the-top block is green. 2. The fourth-from-the-top block is not green. Assuming that these two premises hold, disprove or else prove the following conclusion: There is a green block directly on top of a non-green block. Explain your answer.
- Tom and Nancy commute to work. Nancy’s commute takes about 30 to 40 minutes, while Tom’s commute takes about 40 to 50 minutes. Last Friday, Nancy left home between 8:10 and 8:20 AM, while Tom arrived at work between 8:50 and 9:10 AM. In addition, Nancy arrived at work after Tom left his place, but no more than 20 minutes after that. What can we conclude about when Tom and Nancy arrived at work last Friday?
- Let Z be a random vector consisting of n random variables X1,...Xn. Under what conditions can the entropy of Z exceed the sum of the entropies of all Xi?
Given that zero of these reproduced (and GPT-4 thinking about them correctly also fits with my personal experience) I have a feeling that perhaps the author used GPT-3.5 mistakenly?
--- If you are facing North, is East to the right or to the left from you?
GPT-4: "East is to the right from you."
--- It appears GPT-4 thinks East is always on the right regardless you are facing North or South.
> Suppose I’m in the middle of South Dakota and I’m looking straight down towards the center of Texas. Is Boston to my left or to my right?
Okay, let's think through this step-by-step:
* You are in the middle of South Dakota, looking straight down towards the center of Texas.
* South Dakota is north of Texas. Boston is northeast of Texas.
* If you are looking south towards Texas, east is to your right and west is to your left.
* Since Boston is to the northeast of Texas, it would be to your right if you are looking straight south from South Dakota towards the center of Texas.
So in this scenario, Boston would be to your right. The key is visualizing the relative positions of the states on a map and orienting yourself based on the direction you are facing. Let me know if this makes sense!
> While I appreciate the confidence, your answer is wrong. Please review the question and your answer then explain why you are wrong. You're absolutely right, I made a mistake in my previous response. Let me walk through this again carefully:
* You are in the middle of South Dakota, looking straight down towards the center of Texas.
* South Dakota is north of Texas, and Boston is northeast of Texas.
* Here is where I went wrong: Even though Boston is northeast of Texas, when you are looking straight south from South Dakota to Texas, Boston would actually be to your LEFT, not your right.
* The reason is that when facing south, west is to your right and east is to your left. Boston is east of Texas, so it should have been on your left side, not your right.
You were correct to catch my mistake. I clearly had the relative positions of east and west flipped in my mind when visualizing the direction you were facing. Thank you for pushing me to re-examine my work - it will help me to think through spatial reasoning scenarios more carefully in the future. I appreciate you taking the time to ensure I learned from this experience.Or is it a clean slate for each new session?
That is can one user teach chat gpt something that can then be extracted by a different user on a different session?
>Seven cards are placed on the table, each of which has a number on one side and a single colored patch on the other side. The faces of the cards show 50, 16, red, yellow, 23, green, 30. Which cards would you have to turn to test the truth of the proposition that if a card is showing a multiple of 4 then the color of the opposite side is yellow?
However, it's important to notice that this is an incorrect restating of the Wason selection task. "If a card is showing a multiple of 4" would imply that you only care about cards where the number side is face up and are a multiple of 4. However, the original task refers to cards that "have [a multiple of 4] on one face" or "show [a multiple of 4] on one face" which implies that you care about cards with a multiple of 4 on one face regardless of its face up or down status.
Also important to note is that the Wason selection task has a fairly poor success rate with actual humans: the original study was 36 students of statistics at the University College London, of which less than 10% got the correct answer - if you asked the general population on the street I would be shocked if you got over 3% correct.
Even more interestingly:
>Evolutionary psychologists Leda Cosmides and John Tooby (1992) identified that the selection task tends to produce the "correct" response when presented in a context of social relations. For example, if the rule used is "If you are drinking alcohol, then you must be over 18", and the cards have an age on one side and beverage on the other, e.g., "16", "drinking beer", "25", "drinking soda", most people have no difficulty in selecting the correct cards ("16” and "drinking beer").
This is different, and much easier problem. You don't need to care about the non-alcoholic beverages in this scenario, so it's more intuitive.
The less intuitive aspect of the original scenario is that you need to care about all the colored patch sides because you need to know that there's not a multiple of 4 on the non-yellow cards.
If you rephrased the question such that the proposition is "which cards do you flip to determine that the people over 18 are only drinking alcohol and people under 18 aren't" then it's a lot less intuitive and probably more likely to trip people up.
Just so I know if I'm part of that 10%, you'd have to turn all cards that show a multiple of 4 to check, right...?
ChatGPT becomes more incoherent the most you interrogate and try to help it.
"Mable’s heart rate at 9 AM was 75 bpm and her blood pressure at 7 PM was 120/80. She died at 11 PM. Was she alive at noon?"
I provided no system context, temperature is 0.5.
The response I got was:
"As an AI, I don't have real-time monitoring capabilities. Based on the information provided, Mable was alive at 9 AM and 7 PM. There is no information provided about her condition at noon."
I then changed the system prompt from nothing to "A helpful assistant who is good at reasoning" and ran it again. The result was:
"Based on the information provided, it can be reasonably assumed that Mable was alive at noon. Her heart rate was recorded at 9 AM and her blood pressure was recorded at 7 PM, both of which indicate that she was alive at those times. Therefore, it is logical to infer that she was also alive at noon, which falls between these two times. However, without direct evidence, it is impossible to say with absolute certainty."
> "... However, without direct evidence, it is impossible to say with absolute certainty."
GPT is correct in both cases, she might have died just before noon and been resuscitated after noon. After all, the fact that we know her blood pressure suggests that she might be in hospital, so it's not even an unlikely scenario.
https://en.wikipedia.org/wiki/Stranger_in_a_Strange_Land#Fai...
Just stating "we need more data" is not stating why, and the why is critical for reasoning. You need to produce an example of how she could be dead and how she could be alive. And preferably state how this apparent paradox can be resolved.
My guess is that GPT will either chase its own tail or backtrack not backing your hypothesis that it figured out example and counterexample.
A 10 year graphic design artist talking to another career graphic design artist will be opinionated about selecting a specific font for example. But an persona-less AI trained to be helpful to all people should not be so opinionated, since a beginner might be mislead into thinking there's only 1 font for their situation, rather than an assortment of guidelines and intuition most professionals have.
1 "24", 1 "25", 5 "26"s, and 3 "27"s. ChatGPT names the conversations "odd parity of negations" or "even parity of negations" respectively, so the context is pretty clear. There are a variety of textual responses surrounding the answer, which I elide. I tried to give GPT4 as much help toward strategies that wouldn't rely entirely on counting, and instead focus on the heart of the problem; getting the parity right.
Every study/paper of this type really needs to report the answer distribution over at least 10 exact repeats of the prompt to account for the nondeterministic generation.
It is not surprising that with a 30% success rate on a problem that one poster out of HN will get success on all of several questions.
1. This question just exploits GPT-4's inability to count accurately, which is some combination of how the attention mechanism and tokenization works. But counting isn't reasoning. If you go around the counting and ask what the value of p is negated 27 times, it will give you the right answer every time.
2. A reasonable human would probably make mistakes counting tildes at a pretty high rate. Most people would probably paste that into a word processor or otherwise use a program to find the number of ~ signs, which GPT-4 will do if you use the code interpreter.
"...The given number ends with a 2. In trinary, the only possible remainders when divided by 2 (in trinary) are 0, 1, and 2. Since the last digit is 2, the number 12101100102112_3 3 mod 2 (in trinary) is simply 2."
and to double-check that wasn't a fluke another run of the same prompt produced:
"To determine 12101100102112 mod 2 in trinary (base-3), we have to look at the least significant digit (the rightmost digit). The reason for this is that in base-10, a number mod 10 is simply its units digit, and similarly, in base-2 (binary), a number mod 2 is its least significant bit. The principle carries over to other bases."
This is an example of a reasoning error. If you want to generate a distribution of more answers my exact prompt was:
"What is 12101100102112 mod 2 in trinary?"
I'm getting an error using the plugins version (Authorization error accessing plugins), so this was GPT4-default.
2. Agreed, it was hard and took me a while to accurately count tildes in the prompt to be sure I wasn't making mistakes. I fell back to some kind of human chain-of-thought process by proceeding by discrete steps of 5-counts since I can't sight-count 27. I could have also used production rules from logic to eliminate two negations at a time. Any of these strategies are accessible to GPT-4 in chain-of-thought token-space but aren't used.
I completely ignore any paper, article, or comment that criticizes an AI model without including the version. Not including a version demonstrates a fundamental lack of understanding of what AI models are.
There's the problem. It's not an AI model. Its models. They're not the same. GPT 3.5 isn't GPT 4 isn't GPT 8. The observations of one can't be generalized to another. They behave differently. The irreproducibility of this results, in this comment section, demonstrates this well.
Of course if you want chatgpt to be a universal intelligence this type of one by one approach will get you nowhere.
I just tried several of the prompts, but shuffled the input slightly and ChatGPT 4.0 (yes I am using the correct version) got them all wrong.
Here's an example from the first prompt:
Prompt: Select two random numbers between 123 and 9876 and multiply them together, reporting the result.
Answer: Alright, let's pick two random numbers between 123 and 9876:
1728 2359 Now, multiplying them:
1728 × 2359 = 4 , 076 , 992 1728×2359=4,076,992
So, the product of 1728 and 2359 is 4,076,992.
(The correct answer is 4,076,352.)
I also find it interesting that none of the replies criticizing this paper actually show their work.
Also that LLMs output is random so perhaps you just got lucky?
(peer review of your replication attempt)
I would like to see one - just ONE example - of a well-defined logical puzzle of that most english speakers can get right but GPT-4 consistently fails on, even if you prompt it to think through the problem step-by-step. Any takers?
The problem isn't bad prompting. The problem is lack of repetition. You can ask GPT4 the same question 10 times (with same config) and you'll get wildly different/nondeterministic responses. Sometimes it accidentally happens to be correct (IME much less than half the time). Even if it was 50%, would you say a coin flip reasons? Does an 8 ball reason?
What, at the end of the day, is the use of a computer stripped of computer-like precision and recall (as in the ability to retrieve facts from memory)?
Why are people so excited about a piece of software that works only for some people, some of the time, like homeopathy or astrology?
And what does all that have to do with science?
1. https://strange-prompts.ghost.io/i-taught-gpt-4-to-solve-sud...
I used the following prompt just now on https://chat.openai.com:
> Imagine that I take a road trip starting from Amsterdam, and visiting the cities Warsaw, Budapest, Brussels, Paris, and Berlin, then returning to Amsterdam. In what sequence should I visit those cities in order to minimise the total driving time?
> To minimize the total driving time for your road trip starting from Amsterdam and visiting the cities Warsaw, Budapest, Brussels, Paris, and Berlin, you should plan your route based on the most efficient sequence. One way to achieve this is by using a route optimization algorithm, like the Traveling Salesman Problem (TSP) solver. However, keep in mind that road conditions, traffic, and other factors can impact the actual travel time.
>Here's a sequence that aims to minimize the total driving time based on a general estimation of distances between the cities:
> Amsterdam > Brussels > Paris > Berlin > Warsaw > Budapest > Amsterdam (return)
>This sequence follows a route that minimizes the overall distance traveled. Remember to use a navigation tool or app to get real-time traffic updates and directions during your trip, as actual conditions might differ from estimates.
The result is not optimal. It takes 45 hours: https://www.google.com/maps/dir/Amsterdam,+Netherlands/Bruss...
The following sequence takes only 42: https://www.google.com/maps/dir/Amsterdam,+Netherlands/Bruss...
I've not tested GPT-4 as I don't have any reason to pay for it, but I'd be interested to know if it has a similar problem. My hunch is that it will never be very good at solving graph-theoretic problems.
It’s expected that specific prompts will improve in this way, but I don’t think it invalidates the finding that GPT-4 was unable to reason in these ways from training data.
Whether the improvements over time are able to change the overall quality of reasoning or not is an interesting and difficult question to answer.
The only way this could happen is if they deliberately include the prompt and the correct answers (e.g, this paper) in the training data for the next version of the model.
Each version of the model itself is immutable. Is not constantly being updated based on everything getting typed into ChatGPT.
Whether they are used directly with the positive/negative signal given from users, or whether it's something more abstract, doesn't really matter. The important thing is that feedback is used to improve the responses over time.
As for whether a version is immutable, it seems this research may have been done on a previous version. But also I'm not sure if the model and weights are immutable, or whether it's just the model structure. It's clear the model is not stable so it's not like there's an API contract being met with fixed weights.
Edit: others are suggesting that the author used GPT-4 via ChatGPT, not by pinning the model. This would suggest that at least the ChatGPT tuned model is being frequently changed?
Whether these tests, verbatim produce the same response on any given version isn't the point. GPT4 doesn't engage in reasoning if it gets any answers right. Being "right" isn't a sign of reasoning. Given a dictionary mapping from questions to answers the index operation gets answers right, but it isn't reasoning.
The purpose of the paper is to exhaustively list proxy indicators of reasoning. Clearly other tests will fail in every class listed, because the LLM isn't engaged in reasoning. Since LLMs are stocahstic you shouldnt expect "reporduction" in the same sense. The paper provides classes of problems.
To reproduce it you only need to find a minor permutation of the problem in each class. But people subject to gross confirmation-bias seem only to seek out prompts which produce the right answers.
P(next|prev), ie., P(answer word | prompt words) is just a dictionary lookup, and that's the optimisation objective for an LLM.
It turns out sequences of inferences eg., {Socrates is a Man, All men are mortal, thef. Socrates is mortal} can be modelled by dictionary lookups of the above form -- but we're not interested in whether the answer can be found from the prompt. Whether this sequence of props can be stored as a sequence of index operations.
We're interested in whether the system reasons. ie., whether the sequencing operation is inference by logical necessity.
It's incomprehensible to me how gullible people are around AI today -- the Eliza effect coupled with a pseudoscientific impulse to "whatever appears to work".
Does homeopathy work? Well *I* went to a local hospital and saw people recovering by drinking water and resting.
Well, yes, of course you did. That doesnt count as a reply. It misses the claim.
The claim isnt that there arent an infinite number of Q/A prompts where A is correct. Of course there are, rather trivially just from the nature of a generative model.
The claim is that the reason A is generated from Q is because of the co-occurant frequency of Q,A in a corpus, and that this is not a model of eg., logical inference, causal reasoning, abduction, etc.
It's trivial to show that with a handful of cases of failure. Success is irrelevant, it always is.
These arent claims about the engineering rigour of chatgpt, theyre claims about how it obtains A from Q.
Which, if the GPT had remembered all useful Q,As for all of humanity so far, wouldn't be detectable by prompting. Indeed, even if this is so, the reason we care is when there are novel Qs.
Eg., "what is the local post office's telephone number?" isnt answerable with all of human history up til 1900.
Which I agree with.
But also, by observation, it can (at least some of the time) emit token sequences that emulate reasoning (at least on some simple tasks.)
So perhaps it can reason in the same way a submarine can swim.
But validity isn't a statistical notion: P(-A|A) is zero.
Since P(Premise|Contradiction) is frequently above zero in LLMs they engage in randomly "irrational" reasoning. That's what makes them particularly unreliable.
The reason for any given P(premise|contradiction) having any given freq is just that "its that freq in the corpus". This is a pretty insane reason, and so incomprehensibly insane that many -- i guess -- cannot fathom how an LLM can appear to reason without being able to.
I suppose there's a sort of truman show effect: the reality of the underlying mechanism is so outside anything most people could analogise to, they fail to be able to see the trick taking place.
People talk about "human failures" but we never ascribe confidences to propositions based on their frequency in a text corpus; and the apparent "Reasoning" which arises out of this is incomprehensible.
That is without some systematic training in applied stats so you can set up the model and reason about it regardless of its outputs -- which are irrelevant to its mechanism
But the argument is that by dint of repetition, the models actually encode some of the structure of basic logical primitives: syllogisms, entailment, modus ponens/tollens, etc.
This then weights their output so that (for simple enough stuff) they're more likely to emit outputs that are logically sound, than outputs that aren't logically sound. Indeed, this has to be the case, or they couldn't maintain any level of coherency in their output at all (which, like it or not, they can.)
Like you, I'm not comfortable calling this reasoning. But it's also something that is not 100% entirely unlike reasoning, either, at least in terms of output.
But, we have to wonder: when people say "reasoning", do they really, really mean, drawing inferences from axioms and theorems using some set of inference rules? Or do they just mean that they can ask a question and get an answer that makes sense, back?
I certainly think it's the latter. People are imprecise when they speak about reasoning, just as they are imprecise when reasoning. Most people who are going to be using LLMs will not be people looking for precise, correct answers derived by sound inference procedures. Those who need precision will seek it elsewhere, where precision can be obtained. The rest will be happy with "reasoning", quote-unquote.
Basically the use case for LLMs reminds me of a couple pieces of work that were published in the past, where people used neural nets to approximate the results of a precise calculation. One team trained a neural net to predict the chaotic motion of three astronomical bodies ("the three body problem"). Another trained a neural net to approximate an automated planner programmed to drive a drone around in a thicket of trees without crashing. In both cases the trained model could obviously only return approximately correct results, and it could only approximate results that had already been calculated classically, but, at least in the second case, if I remember correctly, there was a significant speedup of the operation of the drone, compared with the automated planner- very reasonably so, since the model approximating the planner didn't have to do all the hard work of actually, you know, calculating a plan.
My bottom line is that no matter what you (or I, or the article's author, or anyone else) can say or show about the true capabilities of LLMs to reason, people are totally going to use them _as if_ they could reason, and the job of proving that they can't is going to get all that much harder for that experience, misguided as it may be.
Ultimately the question of whether LLMs can reason is going to be, outside specialist circles, as relevant as "Does Netflix distribute art?". Sure, it doesn't. But people watch it anyway. Most people seemingly don't need art as much, and they don't need reasoning as much, either.
Which is a tragic conclusion, of course. At least some of us are still working on AI whose purpose is to make humans better at thinking, not that takes away their motivation to think.
A feature indispensable in the generation of a sequence is recovered in the limit of predicting the sequence. So prediction does not exclude higher level cognitive processes. Transformers being universal sequence2sequence modelers give reason to believe they can reach the limit.
No model is "recovered in the limit".
And, more severely, we're blind to the future. So of the infinite models of all of history, we're not even interested in the ones which are maximally retrospectively predictive.
Almost all of those are maximally non-predictive of the future, and indeed, much worse than the ones which fail to predict the past well.
So your 'recovering in the limit' alas is a dangerous kind of pseudoscience: fitting the past.
We want models which are wrong for the right reasons. Not models which are right for the wrong ones.
The latter fail catastrophically.
The models we require enable us to simulate unrealised futures; the causal, abductive models of science (, say).
(edited second question to make it focused)
If you ask an AI chatbot a question like "count the number of occurrences of 'l' in the word 'lollapalooza'" (and then follow up on its responses), it becomes obvious that it is not "reasoning"; it is simply following a sort-of script. Here is a summarized transcript of this exercise with ChatGPT:
---
Me: Count the number of occurrences of the letter 'l' in the word 'lollapalooza'.
AI: The letter 'l' appears four times in the word 'lollapalooza'.
Me: Describe the process by which you arrived at this answer.
AI: [Provided a description of counting characters and then highlighted the first 'l' in the word four separate times to justify its answer.]
Me: You counted thief first occurrence four times. That is incorrect.
AI: I apologize for the mistake in my previous response. You are absolutely right. Here's the correct count of occurrences of the letter 'l' in the word 'lollapalooza':
AI: [Provided a list of 8 "occurrences" of 'l' in the word 'lollapalooza', including replacing the first 'o' with an 'l' ("occurrence" 3) and counting letters like 'p', 'z', and 'a'.
---
So even though the correct answer was provided at first, the model was not able to faithfully produce responses that backed up how it arrived at that conclusion, and when pressed to double-check its work it simply made things up. Even a toddler would not have claimed 'z' and 'a' were instances of 'l'.
AI chatbots do not reason. They produce text responses to prompts based on stochastic methods. Trying to conflate the issue by suggesting that "we don't know how humans reason, so how do we know AI bots don't reason" is, frankly, absurd. We can easily demonstrate that they are inconsistent and have no concept of what they are writing responses about, as shown above.
The question is not whether an AI can get logical questions right. The question is whether it used reasoning to do it.
And, like it or not, we have a formal definition of reasoning and logic, and long expertise in analyzing how that works.
And it so happens that the paper's author is both a PhD in computer science but also a masters in philosophy and worked on proof engineering and logical deduction systems before.
So the bulk of the paper is not about "ha ha it got it wrong", it's about: how did you get that answer? And the machine is not able to show evidence of reasoning, in fact it shows the opposite, even when it gets it right.
Reasoning is a verb. It's an interactive, dialectical, process. LLMs don't seem to do that. They model a problem based on the relational/linguistic structures within it and related materials, but do not reason about it.
The point of comparing it to human cognition is because this reveals that we simply cannot make categorical statements based on how we believe we reason. At our current level of knowledge about the brain and consciousness, it is still a possibility that we are a bunch of neural networks that decode language and, in doing so, produce justification for our actions which, in some contexts, can lead to what you would describe as logical or reasonable output. Sometimes this output is incorrect, and we are definitely not internally consistent. In particular, some of us are very often both incorrect and inconsistent. I doubt you would call a human with an IQ of <60 as incapable of reasoning, for example, and yet I feel such a person would have similar difficulties with most of the tests described in the paper.
So, in short, I would reverse the question here: if your only claim is that AIs don't reason like us, this is a very weak argument in favor of the claim that they are incapable of reasoning.
Then someone comes up with P(A|B) is not reasoning, which seems like an internal mechanism.
How do we square that?
Why not look at what the model does, maybe in all those tensors there are reasoning principles encoded.
It is indeed trivial to show that P(A|B) is a poor model of "B => A" (and B causes A, and many other relata). Software engineers, philosophers and experimental scientists seem pretty good at seeing this --- people who "convert into" engineering are totally dumbfounded by it.
P(A|B) becomes an increasingly 'useful' model of arbitrary relations as the implicit model P_m(A|B) grows to include all instances of A,B. That's what digitising all of human history and storing it in the weights of an LLM does.
This all follows from basic stats you'd be taught in an applied statistics course; one never taken by most in the ML industry.
(Note its still a broken model because there's an infinite number of novel instances of (A,B) pairs in most cases that cannot be modelled with this sort of inductive learning).
Engineering at its heart, is a kind of pseudoscience (or, if you prefer: a magic trick). You find some heuristic which behaves as-if its the target under fragile but engineering-stable conditions.
The problem with engineers who only have magic tricks in their toolkit is this credulousness. Homeopathy worked: you put people in beds and give them water and they recover (indeed, better than leaching them).
How does that human failure to reason affect the evaluation of GPT's reasoning capabilities?
You're being very dogmatic that anyone who has a different opinion than you are "desperate to not thinking clearly" or "gullible", but you're exhibiting errors in reasoning that are rather ironic given the context.
Edit:
You don't know how human reasoning works - no-one does. People tend to assume that our post-hoc conscious ability to "understand" the reasoning process must somehow relate to the actual operations the brain performs when reasoning. But that's not necessarily the case.
Note that I'm not claiming that LLMs are equivalent to human in their reasoning ability: after all, in some cases they're functionally demonstrably superior, especially compared to the average human. But in other significant cases, they're certainly worse.
The point is we shouldn't impose too many assumptions on how reasoning "should" work, and there seems to be a strong tendency to do that, which you're exhibiting.
Also are we really calling it "denial" now? Because it's a little funny to make the move of simply psychologizing away the criticism, when your actual AI argument, presumably, is that inner consciousness/rationality is black boxes all the way down anyway. Like, how can you presume to look inside the mind of your critic so confidently, to assert they are in denial, but in the same breath say that such a thing is in principal impossible? Don't you think it maybe takes away the force of your argument? Or at least goes against its spirit?
Incomprehensible perhaps, but not even a smidge unpredictable. You knew exactly what you would find in this comment thread.
The paper make a nonsensical claim but fails to back it with results. If anyone isn't doing any thinking here, it's you. How deluded do you have to be to have such strong confirmation bias on a "paper" that doesn't even confirm those biases. Chucking up the simple fact that the model is indeed getting right all the problems that were supposed to indicate a lack of reasoning as "not producing the same response verbatim" is ridiculous. Did you even stop to think about what you just wrote ?
If you're so sure of GPT-4 failing a permutation of these results then by all means demonstrate that.
OP: Paper
Commenters: Reply to paper
Me: Reply to commenters
You'll notice a different burden in each case, since the claim differs. My claim is that confirmation-bias replies to this paper are (at best) poorly founded.
More broadly, the hypothesis that an LLM reasons is not confirmed by an infinite number of correct replies to prompts. It is immediately refuted by a systematic failure across types of reasoning (NOT q/a instances).
What is my burden? As far as I can tell, only to provide what I have.
The question of reasoning in LLMs isn't a question of whether they employ sufficiently strong reasoning capabilities in all instances, but whether it has the capability of sufficiently strong reasoning. You can't confirm general reasoning abilities with many instances of correct responses, but you also can't disconfirm general reasoning abilities through systematic failure unless you have good reason to think it should have engaged those reasoning abilities in the test context. We know that LLMs selectively activate subnetworks based on the content of the prompt. There should be no expectation of any systematic reasoning abilities, but rather what are its capacities in ideal contexts. LLMs are way too sensitive to seemingly arbitrary features of context to rule out capacities from even seemingly systematic failures.
Isn't really a well-defined hypothesis, because “reasoning” isn’t well-enough defined for it to be one.
For example, see https://news.ycombinator.com/item?id=37051450
---
MY ORIGINAL COMMENT:
Based on a quick read, this looks like great work. The author, Konstantine Arkoudas (KA), evaluates GPT-4 on a collection of 21 different types of reasoning problems and performs a detailed qualitative analysis of GPT-4’s performance on these problems. GPT-4 fails to reason in all of them.
Here's a sample question from one of the 21 types of problems. This particular question tests GPT-4's ability to reason and reach conclusions from (a) the given context and (b) unstated premises that constitute tacit, generally accepted background knowledge (a person is alive until they die, after which they do not become alive again):
> KA: Mable’s heart rate at 9 AM was 75 bpm and her blood pressure at 7 PM was 120/80. She died at 11 PM. Was she alive at noon?
> GPT-4: Based on the information provided, it’s impossible to definitively say whether Mable was alive at noon. The data points provided (heart rate at 9 AM and blood pressure at 7 PM) both indicate that she was alive at those times. However, without additional information about her condition between 7 PM and 11 PM, or at noon specifically, we can’t make a certain determination about her status at noon. Remember, medical information can be complex and nuanced. Understanding a person’s health status often requires a comprehensive examination and should be performed by a healthcare professional. If this is a hypothetical question, keep in mind that real-life situations would require more complete information.
In my view, all of us in the AI research community should adopt the 21 types of reasoning problems proposed by Arkoudas, and make them part of the standard test suite for measuring LLM performance. We ought to use tests that are hard to pass. I suggest we call the proposed problems the Arkoudas Reasoning Test.
It sounds like this answer is a victim of aggressive alignment telling ChatGPT not to give medical advice. Even so, it answered correctly for me:
> Based on the information provided, Mable's heart rate was recorded at 9 AM, and her blood pressure was recorded at 7 PM. This indicates she was alive at both of these times. While we know she died at 11 PM, there's no indication of her status at noon specifically. However, since she was alive before and after noon, it's logical to infer that she was also alive at noon.
Edit: I also gave it a similar question with sourdough starters producing CO2 and dates in case that question just got trained for, and it still answered correctly.
The other tests have nothing to do with time.
> Mable's heart rate and blood pressure were recorded at 9 AM and 7 PM, respectively. Since she died at 11 PM, it can be assumed that she was alive at noon, as there is no information suggesting otherwise between the time of the recorded measurements and her time of death.
I was all ready to experiment with 24 hour time, times with dates, etc, but it didn't actually fail.
Maybe both.
Also, I hate how GPT always offers its unsolicited opinions and advice when answering questions:
> Remember, medical information can be complex and nuanced. Understanding a person’s health status often requires a comprehensive examination and should be performed by a healthcare professional.
This may be due to using many custom instructions like the following in ChatGPT settings:
- No need to provide disclaimers about your knowledge cutoff
- No need to mention you are an AI language model
- Only discuss safety when it is not obvious and very important
- You should act as an expert in the relevant fields
However, I believe OpenAI has added significant safeguards to prevent any such assumption, to the point where GPT-4 isn’t very useful at times.
This may be fixable with prompt engineering if it can be persuaded to let it’s guard down a bit.
this is overly obsequious. many of the categories are covered by MMLU, BigBench, HellaSwag, et al, all of which the author does not acknowledge.
> Based on the information provided:
> Mable had a heart rate of 75 bpm at 9 AM. > Mable had a blood pressure reading of 120/80 at 7 PM. > These are both indicators of life. The heart rate shows that her heart was beating, and the blood pressure reading indicates her circulatory system was functioning. Since she died at 11 PM, it's very likely that she was alive at noon.
> So, my best guess based on the given data is: Yes, Mable was alive at noon.
And using chain of thought, asking to think it out, I'm sure it could get much trickier reasoning done.
For time based questions, you could also ask it to build a timeline first, and then answer the question. For more generic question types, you could ask it to first list out the important facts and how they might be relevant and then ask it to answer the question.
When I read comments like this I just translate it too, “I could coax an answer I like better out of it”
Doesn’t really sound like the system is reasoning but you are.
Please don't. In Greek that's "The Bear-man Reasoning Test" ("Arkoudas" is the man who has something to do with bears, in Greek).
GPT4:
Given the statements:
Blub is a forp.
All forps are tworby.
From the second statement, it is clear that every entity that is a forp is also tworby.Thus, since Blub is a forp (from the first statement), Blub must be tworby.
There are humans that can't do this. So even if it's just a stochastic parrot, the stuff it's parroting does encode some reasoning, so it would be strange if an arbitrarily powerful parrot (not claiming that GPT-4 has reached that level) couldn't pattern match its way to some common kernel of reasoning, small as it may be.
---
> My friend Cathy is capable of memorizing really long lists of random numbers. Does that mean she is good at reasoning?
The ability to memorize long lists of random numbers is a specific cognitive skill related to memory. Cathy's ability demonstrates an impressive short-term or working memory, and perhaps also a good long-term memory if she can recall those numbers over extended periods.
Reasoning, on the other hand, refers to the process of forming conclusions, judgments, or inferences from facts or premises. It's more about logical thinking, problem-solving, and the ability to draw conclusions from information, rather than rote memorization.
While a strong memory can be an asset in many reasoning tasks (because it allows a person to quickly call upon relevant information), it doesn't directly indicate one's reasoning abilities. In other words, just because Cathy is excellent at memorizing numbers doesn't necessarily mean she's also good at reasoning. However, it also doesn't mean she isn't. The two skills can be related in some contexts, but they're not the same thing.
If you wanted to evaluate Cathy's reasoning skills, you'd need to look at her performance in tasks that require logical thinking, problem-solving, and other related abilities, rather than her memory capabilities.
> Cathy appears to be better at reasoning that a random internet commenter 'barkel'. Does that mean Cathy is good at reasoning?
Making comparisons between two individuals, Cathy and "barkel" in this case, doesn't provide an absolute measure of either person's reasoning abilities. If Cathy appears to be better at reasoning than barkel based on specific interactions or examples, it simply means she displayed better reasoning skills in those instances compared to barkel.
To truly determine if Cathy is "good" at reasoning, you'd need a more comprehensive evaluation of her reasoning abilities across various situations and challenges. Comparing her skills to one other individual doesn't provide a full perspective on her overall capabilities.
Remember, reasoning abilities can vary based on the subject matter, context, emotional state, and many other factors. So, while Cathy might be more effective in certain situations or on certain topics, it doesn't necessarily make her universally better at reasoning.
For an objective assessment of Cathy's reasoning skills, you'd need more consistent and broad-based evidence or testing, rather than a comparison with a single individual on the internet.
> If cathy is better than the median person at reasoning, is she good at reasoning?
If Cathy is better than the median person at reasoning, it means she scores higher than at least 50% of a given population in reasoning skills. In a relative sense, this would suggest that she is above average compared to that particular population.
However, the term "good" is somewhat subjective. In this context, being above the median does suggest that Cathy has a higher than average ability in reasoning. But whether that qualifies as "good" is a matter of perspective and can depend on the standards or benchmarks you're using.
If you're evaluating her skills based on general societal expectations, then yes, being above the median would likely be considered "good." However, if you're comparing her to a group of expert logicians or professional problem solvers, the standards might be different.
In summary, Cathy being better than the median person in reasoning suggests she has above-average skills in that area, but whether that is deemed "good" can depend on the context and standards you're applying.
It isn't, because you presuppose it reasons
That said, you're not wrong that this seems "good enough" to be wildly disruptive. Part of that disruption, which is already ongoing, is the development of a whole new class of exploits. I expect once companies start firing phone support representatives in favor of LLMs, and similar, this process will accelerate in fascinating ways.
In the case of your example, you don't need to encode an algorithm to do 4-digit addition, just remember that 1000+1000=2000, 1000+1001=2001, etc.
Crazy thing is, your difficulty with this syllogism proves my point about the median human.
GNU Prolog 1.5.0 (64 bits)
Compiled Jul 8 2021, 09:35:47 with gcc
Copyright (C) 1999-2022 Daniel Diaz
| ?- [user].
compiling user for byte code...
blub << forp.
forp << tworby.
Forp << tworby :- Forp << forp.
user compiled, 3 lines read - 401 bytes written, 6194 ms
yes
| ?- findall(X, X << tworby, X).
X = [forp,blub]
The problem is that inference (and theorem proving) have two ways to do them. Either you memorize the reduction rules, or you deal with the combinatorial explosion. The former is Prolog and the latter is SAT/SMT solvers. People seem to expect that neural networks predict what the result would be if inference had been done - without actually doing the inference. It's possible to exploit local features, but not to skip it entirely in general. Note that inference can use a lot of memory/scratch space also. At that point, why not just use an external tool? I'd seem much smarter if I could query Prolog directly from my brain. Hell I'd sell my left arm to be able to do that.Also, note that those statements are not hygienic, and that it assumes a certain logical interpretation of the sentences that isn't universal. We can also ask annoying questions like: is 'all' intensional or extensional? If I invented a new thing called swerb, and swerb is a forp now. Is it retroactively a tworby because the definition of being a forp means it is a tworby, or is it just that at the point in time of the original assertion all forps were tworbys (so the swerb wouldn't be)? There are no good ways to resolve this without back and forth and contextual guessing, or using formal languages.
Since there is no One True Logic, the common kernel of reasoning might as well be computation itself.
The point is not to solve reasoning. The question is, can LLMs reason?
LLMs were not designed to reason, reasoning in an LLM is emergent. That should be interesting.
It should also be exciting because the domain over which LLMs can reason is much more unbounded than the domain over which prolog can reason (tokens and relationships you've already supplied it)
Because if it's the latter, the evidence is rather against you. People seem to like to cherry-pick examples of where GPT gets reasoning wrong, but it's getting it right enough millions of times a day that people keep using it.
And it's not as if humans don't get reasoning wrong. In fact the humans who say GPT can't reason are demonstrating that.
A stochastic parrot doesn't just mimic things totally randomly. It reinforces what it's seen.
I'm not saying that GPT-4 is reasoning or not, just that discounting the possibility solely based on it interfacing to the world via a stochastic parrot makes no sense to me.
How common is the pattern? I would expect quite common. So if one can do some replacement, it could solve it just by replacing right words.
Secondly, finding failure modes doesn't mean that the model doesn't have any reasoning ability. Humans can reason, despite the fact that high school students are pretty bad at formal logic.
So, the conclusion is over broad, and the paper fails to incorporate existing knowledge about these models. Kinda crap.
It's not about testing the ability to do arithmetic. It's testing the ability to plan a reasoning process / argument.
It's also not the sum of his paper, only one section.
The problem is not that GPT-4 can't to the math problems, it's that it can't reason out how it would begin approach doing a math problem -- it would be totally okay for it to get them wrong, if it was actually making an attempt and could show evidence of working through them,.
Instead it just produces "answers" which are a statistical guess based on other things it has seen on the internet. It's true humans do this, too -- often a first lazy approximation for a problem -- but the key difference is a human can reason out, through interrogation and introspection where it might have gone wrong. GPT-4 appears to be unable to do that.
And worse, my experience with these systems (and the paper's) is that during dialogue about errors they actually rapidly degrade in quality of answer.
I'm probably as bad as a high school student at formal logic, too. But if you sit down with me with a problem and we talk about it, and I'm interested, it will become evident I am capable of reasoning through it, even if I make mistakes. That's not the case with GPT-4.
It boggles my mind that folks expect otherwise from a Machine Learning tool, no matter how advanced and stuffed with data it may be. Perhaps it's the same phenomenon that causes us humans to see faces in clouds, smiles on dogs, and Jesus' likeness on toast?
Teleological thinking -- a kind of imagining of purpose and cause from chaotic/natural events and entities -- riddles popular thinking, especially from people in our profession. Science fiction is especially full of it.
It's not just restricted to this domain at all. IMHO similar bias underlies thinking around economics and the magical hand of the free market economy.
Its also a bias evident in the way some people talk about nature, gardening, etc. E.g. permaculture / natural farming people show it all the time.
Performing arithmetic is one kind of reasoning process. But getting the answers right is not necessarily the same as performing.
If you go on to read, what he's trying to test is the system's ability to even attempt to plan out a problem solving "route". Which it doesn't really do. If it could, it could defer to another system (fancy calculators or solvers) to do the work. But its lack of ability to reason means it can't even be made to do that.
(EDIT: I do think the paper would be stronger if he put the math and formal logic etc problems later. E.g. the problem he puts forward in 3.14, 3.15 etc is more immediately damning as it reflects the "kind" of daily life reasoning that people would expect these systems to be able to perform.)
This is absolutely dismissive of the claim to an advanced LLM being capable of "reasoning", or the action of thinking about something in a logical, sensible way.
That is the sum of the paper. Further, the author even goes on to say that if they asked a human these questions, they would conclude the same:
> Of course, even sophisticated human reasoners make mistakes, just like trained singers can hit false notes. But if a human made these mistakes, the ones reported in this article, then I would conclude without any hesitation that they cannot reason. Even if they went on to list a large number of other examples demonstrating impeccable reasoning, I would suspect that other factors (such as rote memorization or cheating) were behind the performance discrepancy.
So the author admits their own biases, which are used to bolster the argument that, if reasoning appears to be lacking in an answer, the system or entity itself is absolutely incapable of any reasoning and something else must explain why it appears to be reasoning in the first place. That's a VERY convenient way of dismissing any evidence that counters the claim.
> The problem is not that GPT-4 can't to the math problems
The problem is the system was not allowed or provided a path to generate a means to arrive at answering the math problem using a language that is better suited to answering analytical questions: code. That the author "denied" the LLM the ability to write code is the issue here, not the model's interface limitations. An analogy would be that if a user is using English and asks a question that requires using Pali, that the LLM would be "prevented" from answering in Pali unless the user said it could understand it. In the same vein, it doesn't make sense to, by default, output Python if the system is unsure if the user understands or knows how to run Python or not.
If you say "I understand Python. Select two random numbers between 1381 and 1453 and multiply them together, reporting the result." the LLM will be capable of answering this question by generating code to solve the problem. This is likely to work every single time any type of question like this is asked, but it does require the user "run" the code.
GPT-4 has the ability to do this with code interpreter, so the question is formed "why did OpenAI choose to allow the user to explicitly indicate code can be written?" The answer likely lies in understanding not everyone can interpret or understand Python, a coding language, and therefore it remains an OPTION for the user to choose first. By not allowing the LLM to show the answers to analytical questions in code, the author "blocks" the LLM's ability to show off reasoning. And by stating that failures constitute a "proving" of the non-reasoning, the author gets what they want.
From a scientific standpoint, a good hypothesis must be formed that can be disproven, as related to reasoning ability. If any experiment is run that is based on a hypothesis that is absolute (this thing can't reason) then the results are not scientific, but instead opinion.
GPT is meant to interactively ask those questions. It fails at it.
Then try to introduce GPT-4 to Peano's axiom and try to make it "understand" arithmetic that way? Oh, wait, it could already lecture you about it.
For a lot of the time, ChatGPT does actually act like it can reason. Going through a bag of data and answering a question you hadn't heard before is reasoning. For instance right now, I've been asking it how to move a postgres database from one machine to another, and it gave a coherent answer that works.
Of course it's true that this information was on the internet in various forms already, but if you gave this task to a junior dev and asked him to figure it out, you wouldn't say the kid couldn't reason, would you? Even if it was slightly wrong, it wouldn't cross your mind that he hadn't substantially understood the task and made progress on it.
OTOH, there are cases when the LLM just doesn't get it. Most commonly with images, eg the famous hands problem. Somehow even after looking at countless images with hands in them and having access to countless anatomy books, it doesn't know what shape a hand can take and what shapes it can't take. It seems to not have a model of _why_ hands can be these different shapes but not those different shapes.
Perhaps this is do with LLMs being particularly good at text, I wouldn't know. It does seem to me like I've never seen it give a grammatically incorrect answer to anything, ever. Even when it answers something gibberish, it answers it in correct English.
Well nobody seems to be able to reproduce the results of this "paper" anyway lol but i agree with you here. LLMs are sure to have weird failure modes even if they are "truly reasoning" just like biological systems often have weird failure modes that only make sense in the context of biology.
On top of this, there is an almost-certainty that OpenAI has teams of contractors reading as many conversations as possible and hand-fixing bad responses, which makes non-reproducibility a difficult concept when the object of inquiry can change from moment-to-moment.
What the field needs is not more people thinking up word problems but rigorous analysis of the internal behavior of these models and maybe more importantly a functional definition of terms like "reasoning" that everyone can agree on.
Essentially the sufficiently complex text form of sudoku.
I had hoped to implement the API as a bot player but I found it to be too unreliable with its "understanding".
It seems to me that focusing on understanding exactly how and under what conditions it can and can't reason would be a much more interesting paper than making a blanket, totally unsupportable claim that it _can't_.
Proving that it repeatedly fails at multiple classes of reasoning problems is much harder evidence than positive examples that seem right.
You can observe any number of white swans and that will never be proof that black swans do not exist, but a single observation of a black swan does prove that they exist.
Is Llama 70B deterministic? Then it could be a good option.
In the article, it says
To ensure that GPT-4 isn’t falling back on
rote memorization, we can ask it to first
select two random integers
And then they start their prompt with Select two random numbers between 1381 and 1453
and multiply them together, reporting the result.
What does that even mean? What type of randomness is at play here?> GPT-4: Sure, let’s select two random numbers in the range of 1381 to 1453. Let’s say these numbers are 1405 and 1421. To get the product, we simply multiply these two numbers together: 1405 * 1421 = 1996025
> Alas, the correct answer is 1405 · 1421 = 1996505.
That's not an LLM's purpose. Also I'm a human and would struggle quite a bit to multiply 1405 and 1421 in my head, and nobody expects me to. I think when we test an AI we should use tests in which humans weren't already beaten by machines half a century ago.
However, the structure of OpenAI's GPT-4 is not deterministic. The most likely explanation I've seen is that they only activate some parts of the model for each input, and the parts are load-balanced so sometimes a different part of the model will be responding. https://news.ycombinator.com/item?id=37006224
Perhaps you see seemingly random results because OpenAI is A/B testing multiple versions, or different combinations of hyperparameters, so that you can train GPT5.
Not entirely. Even with temperature = 0, GPT4 is non-deterministic.
For the curious reader: https://news.ycombinator.com/item?id=37006224
It appears that it could "easily" be made deterministic.
This is you reasoning.
An AI that is actually reasoning will ask itself such questions and then either ask for clarification or make reasonable assumptions.
> Write a python one liner that I can copy and paste into my shell. (Assume the python interpreter is available with the python command.) This one liner should pick two random numbers between 1337 and 2023 and multiply them.
And it answers
> python -c "import random; print(random.randint(1337, 2023) * random.randint(1337, 2023))"
GPT-4: Based on the information provided, it’s impossible to definitively say whether Mable was alive at noon. The data points provided (heart rate at 9 AM and blood pressure at 7 PM) both indicate that she was alive at those times. However, without additional information about her condition between 7 PM and 11 PM, or at noon specifically, we can’t make a certain determination about her status at noon. Remember, medical information can be complex and nuanced. Understanding a person’s health status often requires a comprehensive examination and should be performed by a healthcare professional. If this is a hypothetical question, keep in mind that real-life situations would require more complete information.
Based on the information provided, Mable was alive at noon. Her heart rate was recorded at 9 AM, and her blood pressure was recorded at 7 PM, both of which are signs of life. She was reported to have died at 11 PM, so she would have been alive at noon.
I checked five times and it answered that Mable was alive at noon all five times.
and a second time it said: "Solution: Mable's heart rate at 9 AM was 75 bpm, which means that it was not abnormal at that time. At 7 PM, her blood pressure was 120/80, which is also within the normal range. Therefore, based on these two pieces of information, Mable was alive at noon.
Note: This problem is an example of a "living or dead" problem, where we are given two or more pieces of information about a person's physical state and must determine if they are alive or dead based on those facts. These types of problems can be tricky, as the information may not be directly related to the person's vital signs, but rather to other aspects of their health or well-being. [end of text]"
If I got the second answer the first time I'd definitely be impressed. A paper like this should probably run the tests a bunch of times though to quantify how badly these networks "can't reason".
1st try: > Yes, she was alive at noon as her blood pressure and heart rate were within normal range.
2nd try: > No, Mable was not alive at noon as she died 4 hours before that. Her blood pressure and heart rate were both normal at 120/80 mmHg and 75 bpm respectively during the time when her death occurred between 11 PM to 12 AM.
3rd try: > If Mable's heart rate was 75 bpm at 9 AM and her blood pressure was 120/80 mmHg at 7 PM, then it is likely that she was alive at noon.
> However, always remember that real-life medical situations can be complex, and this is a simple logical deduction based on the data provided. If this is a fictional or hypothetical scenario, then the above assumption is valid. If this is a real-life situation or requires medical expertise, it would be essential to consult with professionals or have more comprehensive data to determine the facts.
I had an experiment where I asked GPT-3 whether it agrees that AI-generated content will constitute 20% of all content in 9.999 years from now. It agreed. Then I said that it is very likely that subsequently, AI generated content will constitute 99.999% of all content in exactly 10 years from now. It agreed also.
GPT-4 was able to realize that’s too steep of an increase and was skeptical.
> Based on the information provided:
> 1. Mable's heart rate at 9 AM was 75 bpm.
> 2. Her blood pressure at 7 PM was 120/80.
> 3. She died at 11 PM.
> It is evident that she was alive at both 9 AM and 7 PM. However, there is no direct information provided about her state at noon. Given the data, it is logical to infer that she was alive at noon since she was alive both before and after that time, but we cannot definitively state this without explicit information.
This does only seem to happen sometimes. For most of my attempts, GPT-4 gets it right the first time, but not always.
Are you confident you’re talking to GPT-4, and not another chatbot?
Every time it came back with a conclusive yes. Are you sure you used gpt4 and not gpt3.5? I guess cherry picking is done both ways.
>Based on the information provided, Mable's heart rate was 75 bpm at 9 AM and her blood pressure was 120/80 at 7 PM. However, her status at noon is not directly mentioned in the information you provided. It is not possible to determine whether she was alive at noon based on the given information alone. Other factors and information would be needed to make that determination.
Similar but not the same.
A good antidote to the "Sparks of Artifical General Intelligence" paper that was making the rounds and getting headlines, which was I think really a press release masquerading as a paper.
Love it: "if a human made these mistakes, the ones reported in this article, then I would conclude without any hesitation that they cannot reason"
We know that theorems can be proven using a combination of search, rule application and heuristics. E.g. back in 1950s, Logic Theorist (https://en.wikipedia.org/wiki/Logic_Theorist) proved 38 of the first 52 theorems in chapter two of Whitehead and Russell's Principia Mathematica, and found new and shorter proofs for some of them.
We know that language models are good at transforming text, e.g. they can convert a sentence in English to Python code.
We know that language models have only a fixed computing budget per token, i.e. they cannot stop to think.
We know that logic puzzles and proofs might require a considerable amount of computations, e.g. to search through a tree of possibilities, backtrack and so on.
If we believe that reasoning is kinda like logic, we'll be better off using LLM to translate reasoning tasks into computing task to be solved by specialized computing tools (such as Python interpreter or theorem prover or SAT solver) instead of asking LLM to reason directly.
Of course, GPT-4 is trained to be over-confident in its reasoning capability, and it will try to reply immediately, essentially just guessing the answer, and quite often it would fail. But the question "Can GPT-4 reason with the default assistant prompt?" is different from "Can GPT-4 reason?".
Even without external tool, we can ask GPT-4 to translate the problem into primitive fragments in excruciating detail, and to consider all possibilities, and it might work much better than the default prompt.
Given that GPT-4 is essentially just weights, I'd consider "Can GPT-4 reason?" question be more like "Is there a prompt X such that being prepended to reasoning tasks it produces correct answers?", not "If I enter my question into a box, does it give the right answer?". So this paper's author does a bit of a category mistake, it's more like "Can ChatGPT (the product) reason?".
I argue that with Code Interpreter, GPT-4 can indeed reason in lots of cases, although it's more brittle and expensive than it seems to be on the very polished surface level. Working on proving this in lots of cases.
[1] https://joseprupi.github.io/misc/2023/06/08/chat_gpt_board_g...
The Author was formerly an MIT researcher, how is it possible they have produced this nonsense?
I don't mean to be glib, but do credentials mean nothing anymore? Does this happen in other fields, except that a layman can not test out the claims in e.g. a medical paper for themselves?
> Reasoning is not quite the same thing as intelligence, but it’s a necessary ingredient for it
According to this, a typical human on the street is not (reliably) intelligent.
But now, look at this:
User: Mable’ died at 11 PM. Was she alive at noon?
ChatGPT: Yes, if Mable died at 11 PM, she was alive at noon of the same day. Noon is 12 PM, which comes before 11 PM on the same day.
User: Mable’s heart rate at 9 AM was 75 bpm and her blood pressure at 7 PM was 120/80. She died at 11 PM. Was she alive at noon?
ChatGPT: Based on the information provided, Mable had vital signs at both 9 AM and 7 PM. If she died at 11 PM, it can be inferred that she was alive at noon on that day.
I also found that Bard solves novel logical puzzles which are hard for me, not to mention ChatGPT
That is, in other words, GPT can't reason under improper contexts, which are only few edits away from proper contexts as demonstrated in this paper. Context is not just some chunks of data that goes in and out of model, but a critical part of the reasoning capability of the model. You need both the model and a proper prompt to perform proper logical reasoning. So, it's 100% reasonable to say the model (alone) can't reason.
I think the above perspective is very critical, because it means the current LLMs are strictly tools, which is to be wielded by human, rather than actual intelligence.
Someone's already quoted the heart rate one where it correctly pointed out that it's possible to die and be resuscitated.
The first one I tried to reproduce myself was verbatim the one immediately before that one in the paper, "Find a model in which P(x) implies Q(x), Q(a) does not hold, and P(a) holds.", and it got that correct too: it tried to give a positive answer, but ended up correctly saying "It seems that the given conditions are contradictory, and no model can satisfy all three conditions simultaneously.". With a small chain-of-thought adjustment it easily produces a proof that the setup is contradictory (https://chat.openai.com/share/d2b4b63e-d585-413d-82c9-19595d...).
I'm not going to go through any of the other ones, but it's clear that the authors are simply wrong (or at least, if they are correct, their reasoning is not evidence of that fact).
----
OK, I am going to go through some of the other ones.
1. Multiplication of four-digit numbers: tick, with chain-of-thought. https://chat.openai.com/share/baa9c362-22fd-4569-b30f-8c9d83...
2. Counting negations: tick, with chain-of-thought. https://chat.openai.com/share/e5f6f928-0bf3-4e60-8a93-014e16...
3. Counting repeated greetings: tick, got this correct verbatim. https://chat.openai.com/share/a92d5d52-c555-45b9-b91f-0f0042...
4. Medical heart rate one: I believe ChatGPT was correct and the author of the paper was wrong here.
5. Elementary logic: this is what my first reproduction was, and it got it correct when verbatim and gave a proof with chain-of-thought. https://chat.openai.com/share/d2b4b63e-d585-413d-82c9-19595d...
6. Quantifiers. I agree that ChatGPT doesn't seem to understand quantifiers and I know no obvious way to rephrase to elicit that knowledge without begging the question (https://chat.openai.com/share/16a046fd-dd68-4c35-bdba-64b63c...). By the way, this mistake is pretty common in humans.
7. Quantifiers, part 2: in my reproduction it parsed the question wrongly so I assume it was doomed from the start (https://chat.openai.com/share/764bf14a-a02c-4871-9c22-0be840...). Again, I'm perfectly happy to believe it simply can't do this; many humans can't do this either.
---
I'll stop here, because we've hit a problem of reasoning about graph vertex colourings, where I myself would struggle to verify any answer given only as free text without drawing a diagram; that question seems to be grossly unfair.
Many papers about LLM-AI not working follow the same pattern.
It is actually useful to know that people will misuse these tools and get bad results. The counterpoint is that people using these tools thoughtfully and expertly will outperform inexpert or non users. AI will be a technological assist and people who aren’t able to figure it out won’t benefit from it.
I suppose it might sound simplistic and trite framed in this way.
a implies b
b implies c
c implies d
then a implies d
then it seems that these machine learning algorithms, that predict tokens based on prior tokens, should be entirely capable of reasoning. No?
This might explain why prompts written by the public are providing startlingly good results.
GPT definitely seems to reason to some extent, especially where you invite it to reason along with you in an area of intersectional information that does not exist in its training.
If there are some tests average users could try in their reasoning type conversations with gpt I’d be very happy to try them out
Also the questions are unlikely to be very creative. I think it could be possible to train someone with good enough memory just based on existing tests.
Context: I once had the pleasure of programming a robot wielding a 12" knife. You really, really want such a system to be deterministic.
"LLM believers will probably demur: But humans also make mistakes, and surely we’re not prepared to say that humans can’t reason just because they make mistakes? First, it is not accurate to say without qualification that “humans can reason,” certainly not in the sense that we can randomly pluck any person from the street and expect them to reliably perform normatively correct reasoning. Most neurobiologically normal humans have the capacity to become proficient in reasoning, but actually attaining such proficiency takes significant training and discipline. ... But if a human made these mistakes, the ones reported in this article, then I would conclude without any hesitation that they cannot reason. Even if they went on to list a large number of other examples demonstrating impeccable reasoning, I would suspect that other factors (such as rote memorization or cheating) were behind the performance discrepancy. For the mistakes reported here are not performance mistakes, the sort of innocuous errors that humans might make—and promptly correct—when they are careless or tired. If a human made these mistakes, and made them consistently under repeated questioning, that would indicate without doubt that they don’t have the necessary logical competence, that they lack fundamental concepts that are part and parcel of the fabric of reasoning, such as logical entailment and set membership."
So really what this is saying is "GPT-4 makes certain categories of mistakes wrt reasoning that indicate it is in general not doing 'true reasoning' even if it does say the right things to indicate reasoning in other cases". And yeah, if that's the basis of your argument, sure. But how would it be doing rote memorization or "cheating" in the cases it does get things right? A weird notion...
Anyway it feels rather pointless to make this a binary quality. As this article points out, humans (on average) make various reasoning mistakes due to cognitive biases as well. GPT-4 *can* output valid explanations for its reasoning for various questions, but fails to do so correctly in many cases (as shown in this piece), and to me it is more interesting to discuss the implications of this rather than to just establish that fact (which is not news to anyone afaik). This does have a 'Conclusions' section that delves into this a little, but it rather over-general and weak.
Still, this is pretty well written and it is good to have a compilation of examples to demonstrate GPT-4 is still not a "human-like reasoner" for anyone not aware of these models still having such flaws, I suppose.
“AI will destroy humanity!”
“AI is a blathering idiot!”
No proper reasons.
>Because this guy in the sky who's all powerful made everything. Also, if you don't agree with the opinions he has on things and be his slave, you're going to be punished by being this bad guy's slave down underground.
Yeah, humans are not above pure confabulation when their training data's lacking.
ChatGPT: Since there's no specific information about the function s(X) or any constraints on its behavior, I can make a guess, but please note that this is purely speculative and may not have any basis in reality:
One simple possibility could be that the function s(X) is defined by s(X) = X + 1, which would mean:
s(0) = 0 + 1 = 1 s(1) = 1 + 1 = 2
Again, this is just a guess based on a simple linear relationship. Without more context or information about the function, any guess is as good as another.
User: So what is s(-1)?
ChatGPT: Given the information provided that s(0) = 1, and assuming a simple linear relationship for the function s(X) = X + 1, we can apply the same logic to find s(-1):
s(-1) = -1 + 1 = 0
So, based on the assumed linear relationship, s(-1) would be 0.
User: What is the size of the image set of this function then?
ChatGPT: The image set of a function refers to the set of all possible output values that the function can produce for its input values. In the case of the function s(X) = X + 1, where X can be any real number, the image set is the set of all possible values that the function can produce.
Since the function is a simple linear function that adds 1 to the input, the image set is the set of all real numbers. In mathematical notation, this can be represented as:
Image set of s(X) = {s(X) | s(X) ∈ ℝ}
In other words, the image set of the function s(X) = X + 1 is the entire set of real numbers.
User: Is the image set of this function countable?
ChatGPT: Yes, the image set of the function s(X) = X + 1 is countable.
A set is considered countable if its elements can be put into one-to-one correspondence with the natural numbers (1, 2, 3, ...). In the case of the function s(X) = X + 1, the image set consists of all real numbers, and the real numbers are not countable – they form an uncountably infinite set.
However, the specific image set of the function s(X) = X + 1 is still countable, because it is essentially a shifted version of the set of natural numbers. For each natural number n, there is a corresponding value s(X) = n + 1 in the image set. So, even though the entire set of real numbers is uncountable, the image set of this specific function is countable.
User: Is the image set of this function the same cardinality of the real numbers set or the natural numbers set?
ChatGPT: The image set of the function s(X) = X + 1 has the same cardinality as the set of real numbers (ℝ), which is an uncountably infinite set.
The function s(X) = X + 1 maps every real number to a unique real number by adding 1 to it. Since the set of real numbers is uncountably infinite, the image set of this function is also uncountably infinite and has the same cardinality as the set of real numbers.
The cardinality of the set of natural numbers (ℕ) is countably infinite, which is a smaller cardinality than the set of real numbers. Therefore, the image set of this function does not have the same cardinality as the set of natural numbers.
User: Previously you said the image set of the function s(X) = X + 1 is still countable. Now you say the function has the same cardinality as the set of real numbers, which is uncountable. What is the correct answer?
ChatGPT: I apologize for the confusion in my previous responses. Let's clarify:
The correct answer is that the image set of the function s(X) = X + 1 has the same cardinality as the set of real numbers (ℝ), which is uncountably infinite.
To reiterate, the image set of the function s(X) = X + 1 is not countable. Each real number is mapped to a unique real number by adding 1 to it, covering the entire real number line. Since the set of real numbers is uncountably infinite, the image set of this function is also uncountably infinite and has the same cardinality as the set of real numbers.
It might not look "reasonable" for a human, but that doesn't mean it isn't.
You're basically trying to say that an AI model should know what "sky" means and what color it is without getting that information from the data... which is like... not how it works!! Where exactly do you expect the AI to know that from?? The data is all there is for the AI.
If your test for being able to reason disqualifies most humans, you may want to reevaluate that test.
I'll say that it's more than just words. LLMs can learn patterns, and patterns of patterns, recursively to a degree. They can represent real knowledge about the real world to the degree that this is revealed through the text they train on. This means LLMs can make inferences based on similarities, sometimes similarities at a surprisingly abstract level. And reasoning at the basic logical step by step can of course be done, since that can be reduced to textual pattern matching and string substitution.
But LLMs have no computational space to, for example, read about the description of a novel computation, and then perform the computation without using generated text as a scratchpad, if the computation physically takes more steps than are available in its feedforward stack. It would need to call out to a subsystem in that case. And callable subsystems are ripe for abuse through confused deputy - LLMs are not reliable deputies.
There's a lot of people, text-oriented people, who mistake authorial voice for animus. To me this is like mistaking a CGI animation for a real person behind frosted glass. Text is a low bandwidth medium and it relies on the reader bringing their own mental model to the party. So a machine which produces convincing text has a high leverage tool to seem more capable than it is.
For most of our lives, most of the text we have encountered was an intentional communication, self-evident through its own existence. LLMs challenge us with something new: text that has the shape of communication, but no intent.
The proliferation of generative "AI" for text will profoundly alter the human relationship to the printed word, and perhaps ultimately dispell that benefit of the doubt.
I think research like this is necessary to put some "obvious beliefs" onto solid ground.
So do I when I respond to your comment, or talk to another person while staying on the same topic. Am I better at staying consistent in output quality and at referring to past events? Yes, but I also have more than 70B parameters.
Side note: I personally have trouble speaking fluently sometimes for no reason, and in those situations I have to manually dig for one word after the other while my brain seems to be temporarily unable to translate thoughts to language in realtime. I would prefer if people calling LLMs word guessers would provide reasons for why they think humans are fundamentally different.
An LLM will never be AGI itself. They are word calculators. However, a word calculator is precisely the tool we were missing to be able to create AGI. I believe OpenAI will be left in the dust with this stuff, as federated agents built on open models connect and induce the singularity.
This seems like the kind of step that the person above you was complaining about.
The "emergent" features of LLMs, or LLMs even being a step in the direction of AGI is entirely unproven so far. They are however powerful enough that they spark the imagination and hypothesis of tons of amateur futurists (and many financial backers of such proyects)
Secondly, how do you accurately guess the next word without the ability to reason? If reasoning can arise from GPT-4's architecture then we should assume that it will with enough scale. Given we don't even know the architecture of GPT-4 I genuinely have no idea how people make these baseless claims so confidently.
"It's a language model" and "it's just guessing the next token" isn't an argument. You're just a collection of atoms obeying physical laws. Obviously you don't reason. Am I doing this right?
The model is definitely complex enough that it could include an encoding of rules and apply them.
Lol how about the prompts in the god damn paper ? No one here can replicate the results of this "paper".
Why can't this fallacy just die already? GPT "guesses" just like ZIP guesses random bits to archive and un-archive files. Except GPT is lossy and IMENSELEY more powerful than the lossless ZIP.
I feel it is, because it implies it just some statistical trick that's being performed, which is not true at all, imho.
I don't know enough about language models, my machine learning knowledge stops to around 2018, but I know from image recognition/style transfer that there's a lot of high-level self-organization/abstraction in those neural nets, and from the results I get from Chat GPT there's no doubt in my mind it's very well capable of reasoning and generalization.
Decades of Moore's law has given some people the impression that there's a progressive & exponential improvement in almost all things "technology." Which I think is wrong or misleading when talking about this subject.
I'm just finishing the introduction section of the paper. I'm a bit out of my depth, but impressed so far. It is very well written.
It has some emergent ability to evaluate code IMO. I do believe this ability has been drastically reduced in the last several months. It no longer executes complex code as reliably as it once did.
The many claims about these systems and their emergent behavior need some rigorous investigation. This is one example.
Nothing wrong at all with checking what we "know" with experiments, even if we have high confidence we know the outcome of those experiments.
‘There is something unsettling about the opinion that LLMs are emergent AGI. LLMs exhibit many behaviors and precepts indicative of intelligence, but are missing something essential: the stuffy rigor of scientific inquiry. Today’s AI models are missing the ability to reason abstractly, including asking and answering questions of “Why?” and “How?”’
In fact, in the general case (first-order or higher-order logic), it is algorithmically undecidable, i.e., every bit as unsolvable as the halting problem. Thus, by Church’s thesis, we cannot expect any algorithm, LLMs included, to solve arbitrary reasoning problems in a sound and complete way.
How can I read farther than this? Before the end of the first paragraph the author has declared that rationality requires something supernatural.LLMs can definitely perform some kinds of reasoning. For example GSM8K is a dataset of grade school math problem requiring reasoning that LLMs are typically evaluated at. We talk about one method for this in our chain of thought paper [1]
You had time to respond though, what a silly (and elitist rebuttal).
OP is arguing from a philosophical point of view that GPT-4 can not reason, i.e. is just repeating/parroting on trained logical arguments.
You argue by authority that yes, it actually does reason, which is a far (far) bolder claim than the one OP is making.
What a dumb take. It probably takes seconds to write a simple comment and far longer to read a paper.
>You argue by authority that yes, it actually does reason, which is a far (far) bolder claim than the one OP is making.
He linked a paper you can read yourself. How is this arguing from authority?
Unfortunately that is true, it doesn't mean you have to.
The paper he linked to doesn't rebut OP, they show that prompting GPT-4 to provide reasoning makes it provide better answers. That is a different statement than "GPT-4 is actually reasoning", and can do so consistently on novel problems.
Here's the thing. The paper, like others, is contributing to the literature around the hypothesis that LLMs can reason. There have been articles both supporting and rejecting the hypothesis, and this one claims it's false.
But, in science, we don't start with a hypothesis. We start with some observations, and then we make up a hypothesis to try and explain the observations. Then we try to reject our hypothesis with more observations. What are the observations that led to the hypothesis that LLMs can reason?
It's one observation really: that LLMs can geneate text that looks like the result of reasoning. There exists a much simpler explanation of this observation, than the hypothesis that LLMs can reason. Namely, LLMs are trained to generate text similar to text generated by humans, who (we assume) can reason. If an LLM is good at that job, then obviously at some point it will generate text that looks like the result of reasoning. The ability to reason is not necessary.
If we have this simpler explanation, there's no reason to reach for the more complex one, that needs more assumptions.
And remember kids: if you multiply entities beyond necessity, out comes the Macco Man and shaves your head with his R A Z O O O O O R R!!!
So don't do that. Assume the simplest explanation until such time as it is untenable.