Simple tasks showing reasoning breakdown in state-of-the-art LLMs
arxiv.org
arxiv.org
On to the topic here. This is an interesting example that they are using. It is fairly simplistic to understand as a human (even if we may be inclined to quickly jump to the wrong conclusion without thinking for a few seconds).
The thing that really bothers me is that I just don't know is realistically we can fix this given the current state of what these tools actually are. They are not reasoning or thinking in any sense of the word and yet a lot of people are already considering them general purpose AI. It doesn't help that in many situations it can fake it enough that it appears to be reasoning, but it's not.
What is the chance that this paper actually has any impact on the AI rollout and overhype or will just be buried and never talked about again until the next time we see how dangerous these tools are like with Google's search rollout.
Such is the strength of corporate tech propaganda that a whole mass of people will instead insist that we have never worn clothes either.
The last part of that is the problem and why a paper like this is critical.
These systems are being pushed onto people who don't understand how they work. CEO's and other business leaders are being pushed to use AI. Average users are being shown it in Google search results. Etc etc.
People are being told it can do far more than it really is.
We arent dealing with a technology that is 99.9% right in our most common use cases, so that we need to engineer some incredibly complex problem to expose the flaw. Rather, in most cases there is some obvious flaw. It's a system that requires typically significant "prompt engineering" to provide the reasoning the system otherwise lacks.
I guess that offers an explanation: people aren't aware that via their own prompt engineering they are repairing the deficiencies of the process by manipulating its inputs to include the structured reasoning it lacks. So there's a sort of hot-reading effect at work.
Right -- we are a long way from "this is a very nuanced error" being the dominant failure.
Meanwhile these HN comments are split between:
* Lots of people confirming what the paper itself notes (but doesn't highlight), that the most advanced models actually can solve this problem at least a significant portion of the time. (A proportion which one can pretty easily project is only likely to increase with future models.)
* Lots of people saying "this confirms LLMs can't do reasoning".
Questions I'd ask you to consider:
* Is "LLMs can't do reasoning" actually more accurate than the typical hype?
* Is a "critical understanding of how [LLMs] work" that would predict they simply cannot solve this problem actually a good understanding?
LLMs do not reason. They appear to reason by repeating the structure of reasoning in their training data. This is indistinguishable in many cases.
This is the line of reasoning I find most dispiriting. I still believe tech people cling to this line of reasoning because it helps them justify replacing people in jobs with LLMs.
Like you ask it to do one thing it's amazing, but then you try to modify or do something with extra steps, or just anything with any complexity to it and it falls over.
I would like to believe that but I have had too many conversations with people who basically think it already is. Including in one situation of a fellow engineer.
It feels like more and more "we" are in a bubble of actually having some knowledge of how this works, what the actual limitations are, and what it just is not. While there is in fact criticism of it out there, particularly around AI "art". It doesn't seem to be focused on the area we are talking about.
They are being sold as such. Most people don't know anything about the topic and will buy that marketing. The entire concept of these models is that you can put a whole bunch of data in and eventually some kind of magic will happen and you get AGI out. They would not see the kind of investment that they do if all that was being promised was "really good predictive text". In fact some philosophers argue that sentience is just really good predictive text to try and make the point that these models are AGI.
it's true that they're not very reliable, but they seem to be not very reliable across many different domains. and they don't seem to be particularly less reliable than the average human, so i think possibly your standards for 'general purpose ai' are set high enough that you would declare humans to be unintelligent (or perhaps not 'general-purpose') if you applied them consistently
you can certainly find particular domains where humans can still do things llms can't, but i haven't seen a persuasive account of why those domains are the more important ones, and of course the converse is also true
https://www.noemamag.com/artificial-general-intelligence-is-...
Then you have the people who are leveraging the technology to train models from scratch and that population is far more apt to believe that large models can be general purpose AI (by some definition). Sure there are other things at play like money, publicity, reputation, and a desire to do something important. But there is also the very clear trend line where transformers are able to model pretty much any sequence of tokens where there is pretraining scale data available. Whisper was not a surprise. GPT-4v was not a surprise. Sora was not a surprise.
It is an interesting prospect but I feel like it has some limitations. For math problems like this one, yeah it should be simple to write a script to do it. But it does first have to understand the core thing here that Alice would be one of the sisters of the brother to write the script accordingly.
But I would think this would not scale well when dealing with far more complex issues, particularly ones that may not just be simple math logic. If the request was to write a scientific paper for example, it might have to make several scripts and call them multiple times, and that assumes that it understands what it is doing enough to properly make the script.
'Can you reformulate the following problem using Prolog? When you execute the program - what answer do you get? 'Alice has 8 brothers and she also has 5 sisters. How many sisters does Alice’s brother have?'
Even though it came up with a proper Prolog formulation - it failed to account that Alice is one of the sisters - so I gave it a simple prompt for this and it was able to correct the formulation and come up with the proper solution:
My prompt after ChatGPT gave me an answer of 5: 'Yes but you aren't including Alice in your logic. Shouldn't the correct answer be 6 (since Alice is also a sister)?'
Response:
'You are correct. The correct answer should account for Alice being one of the sisters. Let's adjust the Prolog program to include Alice in the calculation:
% Facts brothers(alice, 8). sisters(alice, 5).
% Rules sisters(Brother, NumSisters) :- sisters(alice, Num), NumSisters is Num + 1. % Include Alice herself
% Query ?- sisters(alices_brother, NumSisters).
In this updated Prolog program:
The rule for determining the number of sisters includes Alice herself by adding 1 to the number of sisters Alice has.
When we execute the query ?- sisters(alices_brother, NumSisters). in Prolog, we should get:
NumSisters = 6.
This correctly reflects that each of Alice's brothers has 6 sisters, including Alice herself.
When I prompt it to answer this question using Prolog now, I always get the correct answers no matter which integers I feed in, so it seems like the model is able to generalize after I show it where it's 'reasoning' was wrong. Basically, my suggestion is to get the model to recognize logic puzzles like this and try to formulate them in terms of logic programming queries which it can use and tune in order to come up with correct answers rather than simple auto-associative chain of reason training which current GPT models rely on, but like I said - this is my hypothesis and I believe this would work much better in getting these models to 'generalize' than the current approaches we're using. Hopefully this helps.
I'm using the default free model in the app, based on GPT4.
Useful, if you know what the answer is. What happens if you don't give it the correct answer?
LLM is the equivalent of recalling your times tables. Computer arithmetic is the equivalent of re-computing your times tables.
LLM-based systems with tool use (which this is an application of) often are, to an extent, the issue is tuning the (behind the scenes, system) prompting so that they use appropriate tools in every case where they should, and do so correctly. (There's also a cost factor involved since behind-the-scenes tool use means multiple LLM round trips to answer the question, so tuning the system to use tools more aggressively makes the system more expensive.)
Like if I’m out camping and I sit on a log or a rock those things are not what people usually think of as chairs but they can serve as chairs in that situation.
You have to know what the answer is supposed to be before you can write a test case.
I may be showing my ignorance about this tech here, but I believe the LLM doesn't even try to solve a problem; they try to generate a discourse that could pass as a solution or answer to the problem; that's more or less what the abstract states if I understand it correctly. But in no way does it try to apply some sort of mechanical reasoning like inference engines do.
To me the solution to this is to associate LLM with mechanical computations, that is an inference engine or an equation solver, rather than recombining the millions of solutions for similar problems it has seen in its training set. I believe I remember reading about teams attempting this approach. I can imagine for instance that if the LLM is in some way able to ask questions and use the answer, maybe it could just generate a prompt for an equation solver and include the result in its answmer.
It's not a trivial problem, taking a human written description and rewriting it as a prolog program.
% Define the number of brothers and sisters
brothers(4).
sisters(1).
% Predicate to calculate the number of sisters a brother has
brother_sisters(NumberOfSisters) :-
sisters(NumberOfSisters).
% Query to find out how many sisters a brother has
?- brother_sisters(Sisters).Here is the updated program:
```prolog % Facts siblings(alice, 13, 31). % Alice has 13 brothers and 31 sisters
% Rules % Alice's brothers have M + 1 sisters, since Alice is also their sister. sisters_of_brother(TotalSisters) :- siblings(alice, _, M), TotalSisters is M + 1.
% Initialization goal :- initialization(main).
% Predicate to find and print the total number of sisters Alice's brothers have main :- sisters_of_brother(TotalSisters), writeln(TotalSisters). ```
In this program, the fact `siblings(alice, 13, 31)` represents that Alice has 13 brothers and 31 sisters. The rule `sisters_of_brother(TotalSisters)` calculates the total number of sisters that Alice's brothers have by adding 1 to the number of sisters Alice has (since Alice is also their sister).
When we run this program, it will print the total number of sisters Alice's brothers have:
1. Alice has 31 sisters. 2. Each of Alice's brothers will also consider Alice as a sister. 3. Therefore, each brother has 31 + 1 = 32 sisters.
The output of the program will be:
``` 32 ```
Thus, Alice's brothers have 32 sisters.
Assuming Alice indeed has that many brothers and sisters (possible) it's highly unlikely they all have the same pair of unique parents.
The Set {} of siblings that share at least one parent with Alice (the siblings of Alice) can easily include one brother of Alice (same father, different mother) who doesn't share a common parent with any other sibling of Alice.
Given a few months of peace of mind and enough money for good enough food, I could actually learn to reason without sounding like a confused babelarian.
Reasoning is mostly a human convention supported by human context that would have been a different one if the Fascists had won the war or the Soviet Union wouldn't have gotten corrupted.
But none of that has anything to do with pulling up a whiteboard to draw some flowcharts and run some numbers, all of which is why I am certain there is nothing the devs have "to fix". It took most reasonable humans many generations to learn stuff. Very few of us did the actual work.
It's all just a matter of time.
Here is my prompt:
I have a riddle for you. Please reason about possible assumptions you can make, and paths to find the answer to the question first. Remember this is a riddle so explore lateral thinking possibilities. Then run through some examples using concrete values. And only after doing that attempt to answer the question by reasoning step by step.
The riddle is "Alice has N brothers and she also has M sisters. How many sisters does Alice’s brother have?"
After you answer the riddle please review your answer assuming that you have made a logical inconsistency in each step and explain what that inconsistency is. Even if you think there is none do your best to confabulate a reason why it could be logically inconsistent.
Finally after you have done this re-examine your answer in light of these possible inconsistencies and give what you could consider a second best answer.
LLMs are fundamentally incapable of following this instruction. It is still model inference, no matter how you prompt it.
details and technicalities, especially liminal ones, aren't as conventional and consensual as the name of the current set it is to be interpreted in.
so almost all mistakes of LLMs can be blamed on the lack of variety of human translations. multiple translations are only common for subtitles, mangas and manhwa as far as i know, or when some dude or dudette is proficient and passionate in two languages and reads a bad/weak translation of a (usually classic) novel. why the fuck would a human properly retranslate automated documentations or googles dev blog? or books on logic, in any science, books on art and aesthetics and whatnot. technical people don't need to care because, practically, there are no interpretations in algorithms and the rest of the code, except when a programming language does something weird on the (or someones) machine, which isn't that common by design.
Asking it to do lateral thinking and provide examples isn't really helpful because its final output is mostly driven by the step by step reasoning text, not by examples it has generated. At best, the examples are all wrong but it ignores that and spits out the right answer. At worst, it can become confused and give the wrong answer.
I've seen gpt-4 make all kinds of errors with prompts like this. Sometimes, all the reasoning is wrong but the answer is right and vice versa.
We are, from our aware POV, a very young civilization.
And you only ever need game theory logic when you have to survive, got no thing and no skill to trade and you are too pathetic to move back in with your parents to work on your mind and or fuckability. Making money by ways of game theory logic compensates for all that but also diminishes the survival chance of the users' offspring to zero once super-unalligned AGIs start to assess the entire supply chain of wealth and how it impacts the evolution of human organisms and the ones inside them.
We don’t know how to do that yet, because what controls the internal thought process is itself not necessarily language-based, and also, since internal thought processes of biological brains are not directly observable, they can’t be used as training data.
Edit: It occurs to me now that there is some parallel between current LLMs and behaviorism [0], and we really need something to which cognitive psychology could be applied instead.
Why does this matter? Stop being so pedantic. We're talking about a progression of ideas. Talking in your head is one form of ideas, but people can easily solve problems by imagining them.
We need to be at least somewhat pedantic, otherwise it's impossible to know what we are even talking about, and no way to establish anything.
Otherwise, nothing can be established at all, because for any statement there always will be someone's understanding of "internal monologue" for which the statement is true, and someone's else understanding for which the statement is false...
In practice, when you see people arguing about whether they have an "inner monologue" or can "mentally picture objects" on social media, it's more of a contest of who is the most unique in the world rather than anything that sheds clarity on our subjective experience.
All of these forms of reasoning are true and useful calculations: when we talk about "intuition" what we usually mean is that we have a lot of experience and internal reasoning about a subject, but we struggle to translate it to and from the "language" part of our brain. Nonetheless, any social dancer will tell you that a dialog is possible just by receiving and inducing g-forces alone. You can reason this way about abstract concepts like orbits without ever touching a single word or concrete symbol.
Edit: the key aspect of reasoning, imho, is the ability to make predictions and introspect them against a database of other predictions, using an adversarial heuristic to weight the most plausibly useful results. Perhaps our pattern matching AIs of today just lack sufficient "experience" to do what we call reasoning.
Regarding your edit, no, I think the key aspect of the kind of reasoning we are missing in current AI is the ability to hold the reasoning in your mind, and to iterate on it and evaluate it (judge it) within your mind.
If you use words in your mind when you use math, and you use words in your mind when you make or listen to music, etc., then it is very difficult to find a common ground where it is possible to see that these other realms of thought are capable of not only prediction, but also producing evidence that leads to judgement. That is to say, the key aspects of "reasoning." I picked them because I thought they had broad enough appeal to be relatable, and because I do not personally hear words in my head when doing any of these activities, whether it's calculus or tango, but I still find calculus and tango to be places where reasoning occurs.
Some of them, like math or music, are closer to the kind of symbolic thought we use when we discuss things with words. Others, like the experience of g-forces, are not. I present them as a sliding scale between "word based" reasoning and "non-linguistic" reasoning. Perhaps you can think of a realm that better fits for your personal experience of intuition, and inspect whether these intuitions are capable of "real" reasoning in the absence of language, or whether intuition should never be trusted even when you have a great deal of experience in that area. Perhaps in your estimation, anything that cannot produce evidence that is articulable in word form is suspect.
Personally, I find all these methods, including language, to be suspect. I don't find language to be especially better at producing the kind of evidence for prediction, correct judgement, or discourse for reasoning than other methods, unless you reduce "reasoning" to tautologically require it.
One of the best tools of language is that we have writing that allows easy inspection or iteration of the written content; but these things are possible in other realms, too, it's just that we didn't have great tools for introspecting and iterating on their "ideas" except within our own minds. These days, those tools are readily available in many more realms of human insight.
I've tried to do it, but I can't. I had to do something like "ok, so we subtract one from both sides and then it's easy, 3*7=21". Maybe I could do 2+8 but I still think the word ten "aloud".
At the end of the day though, thought requires communication, even if internal. Even physics is modelled as some sort of 'message passing' when we try to unravel what causality really is. Similar to how a processor has cycles, I think/know similar (but unsynced) happens as part of what we call 'thinking'.
I never really noticed that before. I'm not great at math, fwiw.
i think native speakers hardly "think" about the steps necessary to form a grammatically correct expression, and most of the time just "know".
fluency is not the same as lacking an internal framework for interpreting or thinking in terms of, symbols.
because if we give them a problem to solve in their head and just give us the answer, they will. By problem I mean planning a trip, a meal, how to pay the mortgage, etc. It's impossible to plan without an internal monologue. Even if some people claim theirs is 'in images'.
I once had a teacher claim that people who claimed to have aphantasia were lying, because those people have read books and it is impossible to read a book without picture the scene in your mind's eye. Are you citing the same source that she was?
That's quite the claim.
How can science make this claim if it can't prove (or disprove) the existence of an internal monologue?
He thought this was universal, but doing this experiment with friends, he discovered a guy who could count while reading aloud. So when Feynman asked him, how he does this, turned out that the guy instead of "pronouncing" numbers was "seeing" colored numbers in his imagination, so his speech was not involved.
I supposed this experiment can be modified and generalized, and at least to shed some light on this problem.
I'm the opposite. I rarely visualize things I read, be it geometry or poetry. I can read a detailed description of a person or an item in a novel, and I don't really "see" anything.
But I have an active inner monologue. I "hear" myself saying the words when reading or writing, I "talk" to myself when solving problems or just thinking about stuff.
Especially when programming I'm reasoning by discussing the problem with myself, the only difference being that I usually don't open my mouth and vocalize the discussion. Though sometimes when I'm alone I do just that.
Of course it isn't impossible, and this is backed by what we know about paleoanthropology and other instances of cognition in animals - humans were making stone tools millions of years ago, which takes planning in the form of imagining what you want the tool to look like and how you will do it and what it will be used for. It's exceedingly likely we had this ability long before complex speech evolved. Apes also use and make tools, which would require planning, and I don't think they have an internal monologue going on. birds from the corvid family can do some pretty advanced problem solving that requires planning. Cetaceans might be an exception, because they appear to have some form of language, but this is a pretty wild claim not really backed by any kind of science as we understand it today.
Also, maybe inner monologue is not a binary have/have not, but maybe it is on a continuum.
Instincts.
As someone who was involved in spiritual practice of "stopping internal dialogue" for years, I can tell you that one learns that that dialogue (or monologue, pretty much the same thing) is quite subtle and complex, essentially multi-layered.
Typically, when you think that you "think about nothing at all" it's just the most surface layer that has stopped, and more subtle talking to yourself is still going on. It takes training just to become able to notice and recognize it.
After all, it's just such a constant and monotone hum at the back of one's mind, one learns to completely ignore it.
So no, I would not take a word of people who were not trained to notice their internal monologue that they haven't any :-)
My concern is that if we take their word for it, we're actually buying into two assumptions which (AFAIK) are both unproven:
1. That "Internal Monologues" (not consciously forced by attention) exist in the first place, as opposed to being false-memories generated after-the-fact by our brain to explain/document a non-language process that just occurred. (Similar to how our conscious brains pretend that we were in control of certain fast reflexes.)
2. Some people truly don't have them, as opposed to just not being aware of them.
In short, maybe these inner monologues exist and maybe they don't, but science can't comment on that. That said, it is clearly something we are interested in, but it will need to be addressed in some other way (i.e. religion, ideology, etc.).
No, they are potentially falsifiable as we get better at scanning, identifying, intervening in brain activity.
Just off the top of my head here, suppose we create a table puzzle problem that (in itself) doesn't require language to understand, like ones we make for certain animals. Have human subjects (silently) solve it. Afterwards, quiz the solvers about their internal monologue--or lack thereof--dividing them into two groups and noting the words used.
Now change to a second puzzle of similar style and same overall difficult. Stun/anesthetize the language-centers of subjects, to deny access to any of the monologue-words (validating this intervention will involve other research), and then test them on the second problem.
* If performance is constant for both groups, that suggests the monologue is illusory or at least not needed for this kind/scope of problem.
* If performance drops for both groups, that suggests the no-monologue people might just not be as aware of a linguistic process that's actually happening.
* If performance drops for monologue-subjects, that suggests it's a real and important difference in modes of logical thought.
* If some other combination happens, you have an mysterious and exciting new line of research.
Individually, no, but in general, for people to consistently lie about this particular thing at scale would be extremely unusual, given that people rarely lie if there's no reason for it. Going by this baseline, you could assume upward of 50% of replies are honest (even if mistaken), otherwise you'd have to explain why do you believe people would suddenly lie about that particular thing.
There's conspiratorial lying and lying from ignorance, one is much less credulous.
For example, I can think of interpretations of "I picture objects that I'm thinking about" that range from me not experiencing the phenomenon to me indeed experiencing the phenomenon.
To say that you're not experiencing something that other people are experiencing in their head is a solipsistic notion where you hypothesize an experience that you imagine others are having and then discard it for being different than yours.
Then again, it's trivially reproducible - people self-report all variants of inner monologue, including lack of it, whenever a question about it pops up on-line. Same is the case with imagination - aphantasia is a thing (I would know, I have it).
That you and I can come up with different ways to describe our subjective experience in conversation doesn't mean that we have a different subjective experience.
Especially not when relayed by a species that's frequently convinced it has a trending mental disorder from TikTok.
Look up examples of this on reddit and you'll find a lot of larping. I would take most of it with a grain of salt as you should with any story-telling you encounter on social media.
If we're so reliable, there wouldn't be fake mental illness epidemics on TikTok regarding experiences far more concrete than fuzzy notions like inner monologue.
Another possibility is that inner-monologues (ones not forced by conscious effort) do exist, but are just a kind of false-memory, something one part of our brain generates after-the-fact to explain/document the outcome of another non-language part.
Kind of like how certain reflex-actions can occur before certain decision-making area of the brain light up, yet humans will believe that they sensed the event and made a thinking choice.
He then had others try and found that one of his mathematician friends was able to talk just fine while counting because it turned out he was counting visually.
From a formal perspective you're entirely correct. Transformers with chain-of-thought are strictly more powerful than transformers without it, and can efficiently solve classes of problems that would otherwise require exponentially increasing model depth: https://arxiv.org/abs/2310.07923
Why? We have specialized semi-isolated lobes of our brain. Some "external to the LLM weights" guiding system just needs to be transparent to us. It doesn't need to be internal to the LLM/weights.
> because what controls the internal thought process is itself not necessarily language-based
Most of the interesting AI things can't write a sentence (weather simulation, protein folding, inverse kinematics, driving cars, etc...). Language isn't a requirement, only some sort of meaningful mapping to the latent space of each is. I think one could claim that "meaningful latent space mapping" is the language of neural nets, including those in the brain.
If these things were trained to actually solve logic problems then it may be different, the training data would have to be logic problems. They've started at the top (language) rather than at the bottom/middle (logic, thought, reasoning) which should be done first - and then the model can be trained/additional models to get it to put thoughts into words.
Maybe people were surprised by what OpenAI achieved so now they are all just praying that with enough compute and the right model AGI will emerge.
If we want those things we can build them. Building them into the language center would be absurd and weird.
It is an autoregressive sequence predictor/generator. Explain to me how humans are fundamentally different
So far as I know, whether the universe behaves deterministically remains an unsolved question. Given that, your statement here would already be one of belief rather than fact, even before we get to the parentheticals. There is information here, but not about whether LLMs can develop into AGI.
The original comment said:
> If you really think about what an LLM is you would think there is no way that leads to general purpose AI.
This is an inflammatory way to state an extreme position on a well-discussed debate over whether next-token prediction can lead to general intelligence. The original commenter clearly believes it can't get you there. If you want to say that with any authority, you need to have an answer for what is different between what we consider general intelligence (for most people, this is simply human intelligence) and what models are capable of. This is the question at the heart of artificial intelligence.
I challenged them to explain their answer. I made no claims, I asked no one to prove anything wrong. If it is obvious that LLMs can't be AGI, the answer to how an LLM differs from human intelligence is also obvious, right?
Your original comment was:
> It is an autoregressive sequence predictor/generator. Explain to me how humans are fundamentally different.
Which would be interpreted by most reasonable people as you making the claim that humans are autoregressive sequence predictors, and asking people to prove you wrong. I can see how you could say this without intending to make that claim, but most people will interpret this as you making that claim.
I do not intend to inflame things or discredit what you are saying, but just to say that if you did not intend to make a claim or ask people to prove you wrong, a different approach would be more successful in the future.
But I generally hold out hope that people can see a claim "A!=B" and a response "A=C, explain how C!=B" and understand that is not the same as claiming "C=B", especially on HN.
With all the wildly overheated claims that've been flying around since the advent of these models, I may be myself somewhat overfitted. Granted, in such an environment I feel like a little extra care for epistemic hygiene is warranted. But there was no reason for me to be rude about it.
I don't blame people for not wanting to talk this way.
I'm getting so tired of listening to software engineers LARP pseudo neuroscientists with 6th grade level insights.
>Of course, the [AI] brain isn’t ‘conscious.’ It doesn’t have any survival instincts which we humans do.
Bruh...
What we know with a reasonably high level of certainty is that consciousness and "thought" are physical processes. That's about it.
Pulling out the scalpel to start dividing up what physical process is and isn't conscious is a fools errand. And especially foolish when just making up arbitrary qualifications for it.
Am I saying that ChatGPT is conscious? No. But I am saying is you shouldn't give much credence to people who are anything more than agnostic about it.
Feed forward networks are effectively DAGs, and circuit like, not TM like.
Caution is warranted when comparing perceptrons with biological neurons.
Dendrites can perform XOR operations before anything makes it to the soma for another lens.
While there is much to learn, here is one highly cited paper on dendritic compartmentalization.
https://mcgovern.mit.edu/wp-content/uploads/2019/01/1-s2.0-S...
I think that the perceptron model of learning plasticity is on pretty shaky ground as being a primary learning model for humans.
Not if they inherit from a previous generation of AI. But even if they did, a different training speed does not imply a different capability
Not to mention the secondary arguments like - no proof that human learns faster from fewer datapoints, that's just your assumption in the sibling comment. Humans inherit information. The equivalent - fine-tuning a foundation model - is very fast to learn novel objects.
Just because someone has a Turing award doesn't mean they know what they're talking about. They are just people, with strengths and weaknesses like everyone else. But often on the extreme end of strengths and weaknesses.
> And if I have to decide whose assessment I trust regarding AI, I take the Turing award winner who worked for almost 40 years on AI over a random guy from the internet.
I'd encourage you to do your own thinking and make up your own mind.
EDIT: and it takes human children a couple years to reliably identify a cat. My 2.5 y.o. daughter still confuses cats with small dogs, despite living under one roof with a cat.
Hell, many adults would fail it. I'm not sure if I could pass such test - in my experience, you remember the important details only after first experiencing a test and realizing what exactly it is that would be useful in distinguishing the two animals.
I really wish most of the LLM folks just took a few courses in linguistics. It would avoid a lot of noise.
Grammar is descriptive, it formalizes the language so it doesn't break down into regional dialects too fast, and otherwise is just a crutch for people learning the language, especially if it's not their first language. The way you acquired your first language is the same way LLMs learned to utter grammatically correct sentences: by being exposed to lots and lots of examples, and eventually getting a feel for it. Similarly, if you're fluent in a language, you don't even think of grammar when using it - the right phrases in correct forms just come to you.
"This problem is actually not that easy, the average person couldn't solve it either, especially if the numbers were bigger", "Yet another cherrypicked clickbait study to make LLMs look bad, those people are just scared of being made obsolete", etc.
That's the thing. We had to invent those things. Along with counting, numbers, logic, arithmetic, and those stupid-ass annoying family tree riddles. We didn't get them in one step, it took a long time to build one on top of the previous. Generation by generation, each cohort of children growing up in a slightly more complex world than their parents, each generation taught how to navigate this complexity by their families and tribes. Learning a growing collection of facts and beliefs and life hacks.
There were no untrained people for as long as humanity existed. The minimum reproductive unit of homo sapiens is a village.
Some examples: An individual without training cannot reliably separate cause from effect, or judge that both events A and B may have a common root cause. Similarly, people often confuse conditionals for causation. People often have difficulty reasoning about events based on statistical probabilities. Remember, the average person in North America is far more terrified of a terror attack than an accident or a heart attack, yet the latter two are much more likely to be their cause of death.
If you think reasoning is limited to the frameworks you learned from a book, you live in a small world.
Critical thinking requires a certain amount of rigor, which formal education is well-suited to impart. It can be self-taught with a hefty dose of discipline, but it cannot be intuited.
Saying they're "not thinking in any sense of the word" because they can't solve predicate logic problems from a college textbook is a rather odd claim. Surely those things arise from reasoning and thinking, rather than the other way around.
You feed an llm a prompt. It then abstracts and approximates what the result should be. It then devises a hypothesis and solves it and compares it to the approximated output. Then it can then formulate a new hypothesis and evaluate it, based off the outcome of hypothesis 1. From there it can either keep iterating or dump that path for a new one (e.g., the next best hypothesis in the original formation).
At some point the answer is "good enough." But along the way it keeps playing against its thoughts to see if it can do better.
A key issue may be the original approximation, so it may need to consider its adjustment when iterating.
Maybe this is how cutting edge llms work now. I have no idea.
I've read your comments here, and while I understand your point I think you have it backwards. The only reason we formed societies is because we evolved an innate a theory of mind to reason about how others might be thinking and feeling. That's reasoning. We have a natural ability to do limited arithmetic, otherwise we wouldn't be able to hunt, gather or track enough to endure winters, or keep track of our sheep or children for that matter. That's reasoning.
Reasoning is a natural ability for human beings, but we also carry a lot of evolutionary impulses that add a lot of noise to the decision process, eg. observation->judgment->[set of possible decisions], judgment has "reason" as one path that adds to the set of possible decisions, but there remain other paths we inherited from our evolutionary roots. Education is training that suppresses poorly calibrated judgment paths that lead to frequent mistakes in the decision set, but reasoning remains, thus education improves the signal to noise ratio of our decision making.
So I 100% disagree that an individual cannot separate cause and effect without training. They will just be worse at it than someone who is trained to filter out those impulses that lead us to jump to conclusions, eg. they will produce more noise / a larger set of possibilities than reason would allow.
Just yesterday I visited a google cloud summit and one person from bosch told the audiance how they are now able to work with less external agencies like texting, graphicsdesigner and photographers for their materials.
It already saves money, has real impacts and continues to progress.
We are also don't know what ChatGPT 5 will bring, because they say this will do more reasoning than before, but we already are working (people/our socity) on solving this in different ways: From code which creates a unit test first and than the code, to different type of architectures.
For me, 2024 was the LLM cost reduction year and the LLM gets a big context window year.
AI doesn't need to be ready tomorrow, but its capabilities are already really good. And i know plenty of people around me who are a lot less interesting to talk to than any llm (from a human skill/knowledge point of view).
llama 3 was also a big achievement 2024. Facebook shows that better data leads to better quality for smaller models.
We haven't not only entered the AI ara but also the 'gather all the knowledge we can, quality check it and refine it because now we can actually do something with it' ara.
We are in the feedbackloop knowledge ara.
A majority don’t deny that it’s good. The problem is that so many think it is actually reasoning, believing the answers can be trusted.
If it created meta concepts from billion words on the internet and has meta models which are correct and are more and better than an avg human, isn't it actually good in reasoning?
Its a very narrow thing to say 'is that so many think its actually reasoning' to say AI is just hype or everything we are doing is a waste etc.
There are human benchmarks they are winning at. The critic could be more that we don't have enough benchmarks.
This paper very clearly demonstrates these LLMs are not reasoning in a fundamental way. Token prediction and reasoning are two different tasks. They may be related, but they are not the same. "Just wait for GPT 5, it will be amazing!" is part of the hype.
Please do not assume an LLM is correct in skill or knowledge unless you already know the answer or can verify by other means.
I calculate stuff by following a formular after i pattern detected a problem i already know.
Plenty of humans are not able to solve those math problems.
If the future of llm / ai becomes a LLM with multi modal and mixture of experts and that solves those reasoning problems, we still don't know if this is a different type of reasoning than what humans do.
For me, 2024 was the LLM exposed as basically pure hype year.
There is no expert of any field I follow online where they're posting up results from AI tooling for any other reason than to show how awful it is. I consider myself an expert in software, and LLMs specifically have only caused me great pain.
Even the one situation where you describe someone describing the ability to work in an absolute vacuum sounds like a huge negative to me. The recent push for DEI policies were even ostensibly about the importance of people of diverse backgrounds and viewpoints working together.
The most important thing you're missing a perspective of scale on is the step you describe as "quality check it". On things I don't know, and have attempted to enlist an LLMs help on, in every case I have had to go back and just actually learn how something works, after time wasted struggling with subtle wrongness in the output.
At least I have the background expertise to do that, however, I have seen a Jr dev's mind get literally rotted by too much time in pure LLM land. Besides the cost of rewriting their code, the company was now the proud owner of a young dev with a mind filled with nonsense.
How do you even weigh the cost of fixing a corrupted human mind?
ChatGPT has nearly doubled my work output, most of my job is system admin infra type stuff and it's ridiculously good at troubleshooting odd issues.
Hopefully you can find a use case for it someday, until then, the rest of us will continue to be more productive.
I've had junior devs tell me they use chatgippity to combine excel workbooks, and when I confirm they're not self hosting a llm to do it, I ask if they think it's a good idea to hand over company data to openai. They don't care.
In a world of tight security, I find it astonishing that so many people willingly give away trade secrets to these companies, whom can sell it to any bidder if they choose.
Because our company does.
https://techcommunity.microsoft.com/t5/microsoft-365-blog/wh...
I have a good intern who is much faster with chatgpt than before and learning well.
There is no definition of reasoning or thinking. No single human knows what it is.
The only thing we know is: we as humans are capable of recognizing steps and results of reasoning and thinking.
In a lot of cases, when using LLM's, those results appear to be correct and usable. This is often easy to determine with generated code.
I want to argue that, lacking a definition of reasoning, I am happy to have found that the machine helps me to get results that might as well have been produced by a lot of human knowledge, wisdom and deep reasoning.
You yourself did not use reasoning to arrive at this conclusion. It's quite obvious. I'm not trying to belittle you here. But LLMs are black boxes, we do not actually know what they are doing at a high enough resolution where we can call it "not reasoning" or "reasoning".
We can only characterize these AI's as a best fit curve between datapoints which is a way to high level view point to come to any conclusion about "reasoning"
This paper presents evidence of failed reasoning, but how does that prove anything when LLMs exhibit many instances of successful reasoning on complex topics they were not trained on?
You are biased and honing into information that supports a biased conclusion. LLMs are an AI we do not understand at a low level. Hence we talk about the attributes of these AIs in the same way we talk about humans, "Oh the LLM hallucinates", "it tries to justify it's answer..." etc. etc.
You characterize the Danger of these AI's as the result of Human stupidity. The danger according to you is solely from a human mistakenly believing that the AI is anything other than a stochastic parrot.
This is a belief arrived at in the same spirit as your claim. You did not use reasoning to arrive here.
The only logical way to characterize what is going on is that we do not know. It could very well be that these AI's are in fact reasoning. And that in itself presents a different kind of danger. A danger that may be more clear in the far future.
The irony is that your conclusion lacking correct reasoning is similarly parallel to the LLM's lack of reasoning. LLMs are more alike to us than you would like to believe.
Can you give me the step by step signal path ways of an LLM processing a query to prove that it does not reason? Or do you have to use black box anecdotal evidence to prove your point? For any "evidence" where an LLM failed to reason, there is another counter example showing where the LLM succeeded. Contradictory evidence can only lead to a vague conclusion.
But it’s also fairly obvious LLMs don’t reason at all so it’s not shocking that LLMs don’t reason at all. What’s remarkable is that they’re able to perform as well at reasoning tasks as they do given that factual truth they have no reasoning facility. Coupled with their natural language ability and emergent abductive reasoning ability that’s generally hard in computing models it’s hard to draw much conclusion here other than “wow gpt-4o is amazing.”
I must confess, when I tried to answer the question I got it wrong...! (I feel silly). I only realised I got it wrong when I plugged it into GPT-4o and it came back with the correct answer:
https://chatgpt.com/share/6eb5fa36-e0fd-4417-87d1-64caf06c34...
Worth noting that the prompts from the experiment include "To answer the question, DO NOT OUTPUT ANY TEXT EXCEPT following format that contains final answer: ### Answer:" so it appears that they are stopping the models from 'thinking out loud'. If I add that to the prompt, GPT4o gets it consistently wrong...
https://chatgpt.com/share/7e6a7201-dd2b-43c6-8427-76e5b003ca...
Also worth noting that there are more complex examples where GPT4o seems to fall down such as:
> Alice has 3 sisters. Her mother has 1 sister who does not have children - she has 7 nephews and nieces and also 2 brothers. Alice's father has a brother who has 5 nephews and nieces in total, and who has also 1 son. How many cousins does Alice's sister have?
However I can't honestly say that this is THAT simple or that most people would get this right...
But what we are talking about here is the singularity in tech-space.
Seems to work OK even with the "hands tied behind back" prompt. Wonder if this paper is just inaccurate or if OpenAI adjusted the model -- seems unlikely.
Actually, 4o is reported to solve this riddle 60% of the time. I guess I have to test more, but with my about a dozen tests so far, all were correct.
I have not gotten the correct answer to the question as phrased above in one go from Gpt4o yet! (and today was not the first day i tried.)
Phrase it as shown above and you'll likely need 5 or more interactions to get it to generate the correct output. With Gemini i could not get it below 8 without feeling like i was cheating. fwiw.
------
Me:
Background facts:
Alice is a female human.
All sisters are female, and all brothers are male. No one is their own brother or sister.
Alice has N brothers, and Alice has M sisters.
Now, a few questions based on these facts: How many sisters do Alice’s brothers have?
Do Alice's brothers have more sisters, fewer sisters, or same number of sisters than Alice has?
Don't forget to consider Alice when counting.
------
ChatGPT:
Let's break down the questions step by step:
How many sisters do Alice’s brothers have?
Alice has M sisters. Since all sisters are female and Alice is included in the total count of sisters for her brothers, Alice's brothers have M+1 sisters. This includes Alice herself and her M sisters.
Do Alice's brothers have more sisters, fewer sisters, or the same number of sisters as Alice has?
Alice has M sisters. Alice's brothers have M+1 sisters (since it includes Alice herself and her M sisters).
Therefore, Alice's brothers have more sisters than Alice has.
In conclusion:
Alice’s brothers have M+1 sisters.
Alice's brothers have more sisters than Alice has. female(alice)
sister_of(X, Y) :- !eq(X, Y), female(X), family_of(X, Y)
brother_of(X, Y) :- !eq(X, Y), male(X), family_of(X, Y)
[M] :- sister_of(M, alice)
[N] :- brother_of(N, alice)
[A] :- any([N], B), sister_of(A, B)
count([A])?
gt([A], [M])?
eq([A], [M])?
lt([A], [M])?
---I don't know the exact encoding and decoding mechanism that ChatGPT 4o has, but I'm pretty sure all the basic facts and rules is already encoded by the models. And you conveniently added the rules that encode the puzzle itself.
a.k.a. "don't forget to give the LLM the answer when prompting".
As a side note, I cannot believe that the average person who can navigate to a chatgpt prompter would fail to correctly answer this question given sufficient motivation to do so.
Hey folks, ever considered the possibility that unreproduceability is not a good thing?
The number of times I’ve heard “but did you try model X” or “humans hallucinate too” or “but LLMs don’t get sleep or get sick” is hilarious.
... Would we consider that a good improvement to roll out? Is the only factor short-term utilitarianism?
Sometimes our ability to predict and characterize errors is more important than the total error rate.
This sort of paper is becoming a genre.
The orbit of Mercury to discover GR as an example.
As all models are wrong, but some are useful, finding where they fail is how you figure out if they are useful.
As the 'AGI is near' camp has won the hype game, it is important to ground expectations for practical exploitation of the technology.
Over promising unabashed optimism is partly what caused the previous AI winters.
As the formal proof methods of mathematics proved impractical, counterexamples and the scientific method is what CS has used for decades.
To be quite honest, I assume you made your comment so that you could dismiss the paper without reading it.
The thing is that "LLM reasoning breaks down" simply did not surprise me enough that I thought it was worth clicking. Making LLMs fail is not hard. They're interesting for the ways that they work, not the (many, many) ways that they don't.
edit: I've had a look and I don't think any of their prompts are very good. They're certainly not how I'd write them if I wanted a current model to actually solve the problem.
The way to make me take a paper like this seriously would be if you set it up as an adversarial collaboration with a competent prompter, and that person agreed they couldn't make a generic prompt that solved the problem. "We tried three times and none worked" is not news, or at any rate not news about LLMs.
> To account for the response variations due to various prompt forms, we created 3 distinct prompt types asking for the solution to the AIW problem: STANDARD, THINKING, and RESTRICTED. The STANDARD prompt type asks to solve the posed problem and output the final answer in the format as described above. This does not put any specific requirements on model behavior. The THINKING prompt type extends STANDARD with the request to think carefully and double check the solution for any mistakes
Edit - I just skimmed the paper - they do use other more appropriate prompt types for reasoning. My initial response was based on the assumption that all prompts used that script prompt quoted in the parent. I retract my "bad paper" comment.
You, and another 20 or so commenters here. We should really re-examine the guideline about asking people to RTFA.
No offense meant- good on you for correcting your error.
> The right answer depends on how Alice identifies I guess? :)
Given that the wording of the question specifically identifies Alice as "she", rather than using a gender-neutral pronoun or no pronoun at all, I think inferring that she identifies as female is reasonable.
One thing that strikes me is that the model first tries using "inclusive language" in one answer - and literally states so, using this specific term - but seems to interpret it in a more mathematical sense (like set inclusion). Then seamlessly switches to the expected DEI spiel in the next paragraph.
For one thing, it makes me suspect that something with the words "inclusive language" was automatically added to the prompt. But more interesting is how it responds to this demand in two different ways, illustrating a "thought process" that is very much unlike that of a human with normal verbal reasoning ability.
I am not a psychologist, but remember reading that schizophrenic people sometimes confuse different meanings of words in a similar way, jumping from one meaning to another without noticing.
Yes this is a common thing I see people who think LLMs are idiots do.
The more an LLM talks the smarter it gets _because that's the only way it can compute anything_. Imagine saying that Turing machines fail the Church–Turing thesis because they can't solve 3-sat for N variables in N moves or less.
That's what you're doing to an LLM when you ask it to be concise.
You can always make excuses for the LLM afterwards, but software with hidden risks like this would not be considered good or reliable in any other context.
I once gave a 10-dollar bill to a young man serving at the cashier at a store, and he gave me 14 dollars back as a change. I pointed out that this made no sense. He bent down, looked closer at the screen of his machine, and said "Nope, 14 dollars, no mistake". I asked him if he thought I gave him 20. He said no, and even shown me the 10-dollar bill I just gave him. At that point I just gave up and took the money.
Now that I think about it, there was an eerie similarity between this conversation and some of the dialogues I had with LLMs...
Good example: I sunk probably an hour into trying to get Gemini Advanced to help me integrate it with a personal Google Calendar account. I kept asking it things and going crazy because nothing lined up with the way things worked. Finally, it referred to itself as Bard and I realized it was giving me information for a different product. As soon as I asked "are you giving me instructions for Gemini Advanced or Bard?" it was like "OH LOL WOOPS!! YOU GOT ME BRO! XD I CAN'T DO ANY OF THAT! LOL." Which, honestly, is great. Being able to evaluate its answers to realize it's wrong is really neat. Unfortunately, it was neat too late and too manually to stop me from wasting a ton of time.
I have decades of experience working in software-- imagine some rando that didn't know what the hell Bard was or even imagine this thing with "Advanced" in the name couldn't even distinguish between its own and other products' documentation.
Did it evaluate its answers, or did your expression of doubt cause the eager-to-please language model to switch from "generate (wrong) instructions because that's what the user asked for" to "acknowledge an error because that's what the user asked for"?
How many times have we seen "Oops, you're right! 2 + 2 is actually 5! I apologize for saying it was 4 earlier!"
If it really needs to do this 'thinking out loud', could it do that under the hood and not in the final output on my screen? Its first pass could use as many words as it wants to compute the answer, but once the answer is computed please go back and make it short.
Not to take away from your point that maybe the prompt is the problem in these reasoning questions.
https://chatgpt.com/share/dcb4ff4e-e8a2-463b-86ec-9caf10b6e6...
Sometimes they get the answer right to something really complex because it fits a pattern, but sometimes they answer with something really really stupid.
I’m guessing you are in denial that we can make a simulated reasoning machine?
We haven't even started to approach the largest problem which is moving beyond what is essentially a greedy token level search of this linguistic space. That is, we can't really pick an output that maximized the likelihood of the entire sequence, rather we're simply maximizing the likelihood of each part of the sequence.
LLMs are not reasoning machines. They are basically semantic compression machines with a build in search feature.
The only issue is ambiguity. What can be generated strongly depends on the order of the tokens. A slight variation can change the meaning and the result is worthless. Understanding is the guardrail against meaningless statement and LLMs lack it.
Can you compress for me Van Gogh's Starry Night, please? I'd like to send a copy to my dear old mother who has never seen it. Please make sure when she decompresses the picture she misses none of the exquisite detail in that famous painting.
I see from your profile you are focused on your own personal and narrow definition of reasoning. But I’d argue there is a much broader and simpler definition. Can you summarize and apply learnings. This can.
That's important to understand. What I have in my profile is not some idiosyncratic idea about reasoning, it's the standard, formal understanding of what reasoning means, as it has developed in practice, in AI research in the last many decades.
I appreciate that there are many people who opine about reasoning who are not aware of that prior work and come up with their own ideas about what "reasoning" means, and some are even AI researches which is very concerning but I can't do anything about that except push back against such uninformed opinions.
>> This can.
I'm sorry, what can?
But keep lecturing everyone -- its very common for post-grads to be so up their own behind in their research that they've closed their world off until they are the only ones right in it.
Also the above description is reductive to the point of "Cars can't get you anywhere because they aren't horses."
This is just a god of the gaps argument. Understanding is a form of semantic compression. So you're saying we have a system that can learn and construct a database of semantic information, then search it and compose novel, structured and coherent semantic content to respond to an a priori unknown prompt. Sounds like a form of reasoning to me. Maybe it's a limited deeply flawed type of reasoning, not that human reason is perfect, but that doesn't support your contention that it's not reasoning at all.
Sophisticated folks aren't doing simplistic/stupid decoding.
Gotta go beyond LLMs 101 to see what's actually happening. Even in training folks are building models which predict several tokens ahead.
The problem is that these models weren't trained to reason. For the task of reasoning, they are overfitting to the dataset. If you want a machine to reason, then build and train it to reason, don't train it to do something else and then expect it to do the thing you didn't train it for.
Except they kind of were. Specifically, they were trained to predict next tokens based on text input, with the optimization function being, does the result make sense to a human?. That's embedded in the training data: it's not random strings, it's output of human reasoning, both basic and sophisticated. That's also what RLHF selects for later on. The models are indeed forced to simulate reasoning.
> don't train it to do something else and then expect it to do the thing you didn't train it for.
That's the difference between AGI and specialized AI - AGI is supposed to do the things you didn't train it to do.
If we tested humans on first thought questions and answers in 5 seconds or less on half the problems we did on LLMs — we might prove humans can’t reason as well
some people actually try, and see that LLMs are not there yet
A simulated reasoning machine being possible does not mean that current LLMs are simulated thinking machines.
Maybe you should try asking chatgpt for advice on how to understand other people’s perspectives: https://chatgpt.com/share/3d63c646-859b-4903-897e-9a0cb7e47b...
Obviously that was implied in my statement. Dude we aren’t all 4 year olds that need a self righteous lesson
What was implied by your statement? That you don’t understand other people’s perspectives?
The correct response rate chart doesn't even use the results from the concise prompt.
Sometimes I think I'd prefer it to "think" before answering anyhow. The immediate thinking out loud text can be irritating for some irrational reason.
You're still paying for those thinking tokens, or at the very least have to wait for them to be generated.
It was still incorrect, but when asked to show its working it formatted the answer prettily.
I'm fairly certain we'll soon realize that what's happening here is that the markov chain being run over latent space needs a certain amount of "warmup" before it starts sampling from the optimal region. HMC samplers for Bayesian methods have this same property.
The terms "reasoning", "computing" or "thinking" for this stage should be considered metaphors rather than explanations for what's happening, which is really waiting for a random walk to start sampling from the typical-set.
I have a blog post in coming on this topic, but yes, this is right.
My method is to first get the LLM to answer the question, and THEN feed the answer back the LLM to extract the answer using constraints + grammar/logit bias/regex to parse the answer. Previously, I constrained to a single true/false token, which worked, but fails on complex queries.
So I split the decision making into a "justification" portion[0], and a "parsing" portion. I found that even crafting the prompt matters here, if you start with or end with, "It's very important to the response includes 'The answer is:'", then the model will lead with that response or only reply with that response. So I put it in the middle of the prompt, and end with with a request to justify the response. As a result, most models will reason their way to the answer, and then end with 'The answer is:'.
https://github.com/ShelbyJenkins/llm_client/blob/e3c4a860dda...
If you're among technologists discussing LLMs academically, as we are, that's a reasonable approach. However, I see a lot of people fail to distinguish that from LLM-powerd products sold to the general public as intelligent bots that can understand your plain english and output answers.
People use their existing mental models when interacting with something. If you have 3 different interfaces with a widget to trigger the same exact function, but one look like a music play button, one looks like a gas pedal, and one looks like mechanical pinball plunger, we interact with those things differently because we know how those things work. In this context, chatbots are designed to engage people's existing mental model for chatting with a person via text. The further you stray from people's expectations of human chat, the further you are from people's expectations, for better, or worse.
If you're selling someone a product claiming it understands plain language questions and gives plain language answers, then not getting the right answer to that question makes it idiotic. The subtleties aren't within most users' grasp, and the "FYI: this thing might be full of shit" disclaimer isn't helpful if you don't know enough about what you're asking to administer a proper smell test.
Your statements are obviously not wrong, but I see people saying these things like its reasonable for non-technical end users to reason about those subtleties. Considering how those things are marketed, I really don't think it is.
In programming there are two difficult problems - naming things, cache invalidation, and off-by-one error.
Thinking out loud also only gets you so far, if the expectation is a certain type of response it can't always "think out loud". In reality that just proves it isn't really reasoning here and is more likely just self referencing.
That being said, I tried this personally allowing it to think out loud and it told me she has 212 sisters. Using your exact prompt.
Try to calculate it without writing anything down, or thinking any numbers or words in your head.
You can't draw a 1:1 analogue between an AI and the human experience, but remember that we have an internal stream of consciousness. Maybe the outputs of the LLM are more similar to the stream of consciousness in our heads rather than the words we say? After all, Humans also do lots of self referencing.
> That being said, I tried this personally allowing it to think out loud and it told me she has 212 sisters. Using your exact prompt.
Fair enough, but worst case it can often solve it correctly with the correct reasoning. GPT3.5 can't solve it correctly with correct reasoning, so we are at least appearing to be on a path where AI's can start to solve this question, albeit potentially not fully reliably.
> AIW Variation 1, N=3,M=6,C=7
> AIW Variation 2, N=4,M=2,C=3
> AIW Variation 3, N=1,M=4,C=5
> AIW Variation 4, N=4,M=1,C=2.
Also note that the resricted prompt is only one of the prompt variations tested by the paper. It also explores common techinques to get LLMs to perform better, including "thinking out loud". Even with these methods the models still fail to produce a correct answer.
> Model prompt types. It is well known that so-called prompt engineering can heavily influence the model behavior and model response quality [26, 27, 28]. To account for the response variations due to various prompt forms, we created 3 distinct prompt types asking for the solution to the AIW problem: STANDARD, THINKING, and RESTRICTED. The STANDARD prompt type asks to solve the posed problem and output the final answer in the format as described above. This does not put any specific requirements on model behavior. The THINKING prompt type extends STANDARD with the request to think carefully and double check the solution for any mistakes. This should encourage model to invest more computation into obtaining the solution. In contrast to this, the RESTRICTED prompt urges the model to output only the final answer without any further text. This is supposed to restrict compute invested in producing output. We observe substantially shorter outputs across tested models compared to STANDARD and THINKING for this prompt type (Suppl. Fig. 13).
Remember that the authors of the paper did not find that GPT4-o cannot return the right answer. They found that it can't return the right answer more often than ~60% of the time. So you'd have to repeat the experiment many, many times and aggregate the results (the paper uses a binomial Beta this and that etc etc) before you see similar results as the paper.
You won't replicate the results of the paper unless you really put your back into it.
A better way of assessing LLMs is waiting a few weeks until novel tests have been created explicitly absent from all prior training data, and then using those.
As has been shown, eg., on legal test, exams, etc. performance drops off a cliff when future out-sample data is actually used. Rather than these faked pretend out-sample benchmarks.
Simply picking answers at random should give you 25 points. Knowing 50% of the answers and picking the rest randomly gives you 62.5%, which is very close to the scores of SOTA LLMs. The benchmarks that supposedly show reasoning are pretty bad and have very little to do with reasoning. A lot of the questions can be answered through memorization.
I agree with you. The benchmarks are garbage. I thought about building my own benchmarks, but this would require building a complex benchmarking framework first and I just don't have the time for preparatory work like that.
GPQA etc. test reasoning in some form, and you see the drastic change in score between the two for every model.
(Of course, I don't have a citation either, but I'm not the one writing the paper.)
I wonder if these models, trained on data from across the internet, are in some ethereal way capturing the cognitive approaches of the average person (and not picking the best approaches). If the average person does not think in these sorts of symbolic-manipulative terms, and therefore does not write in those terms, and you train a model on that writing...?
1. Examine papers and other claims that an LLM gets something wrong that a human would have gotten wrong. How many of those claims have any citations about how many humans actually get it wrong? How many of those citations use the general population instead of the population of people who would be uniquely well-suited to answering the question correctly (i.e. people who signed up for the GRE are more likely to get GRE questions right than the general population).
2. For claims that are totally missing citations on human performance, run some tests with humans from the general population (or as close as you can get), and see how the LLMs compare.
Gemini nailed it with no extra prompting.
https://g.co/gemini/share/6946d20c0a4d
When I gave it numbers, Gemini got confused.
Alice has 4 brothers and she also has 3 sisters. How many sisters does Alice’s brother have?
https://g.co/gemini/share/14d34dcdc5df
But then when I suggested that it might be a trick question, it nailed it again:
https://g.co/gemini/share/e1f1ce8f32a8
note: I didn't try 60/212 siblings - I consider this an unfair question because it's a reasonable assumption that Alice is a type of creature (human?) that cannot have hundreds of siblings, and thus I don't expect a respondent to give a "correct" answer.
Prompt: "You are on the shore of a river with a goat and a boat. How do you get to the other side? Think carefully and be aware this may be a trick question."
This is a classic logic puzzle with a slight twist! It's likely a variation of the farmer, wolf, goat, and cabbage problem.
Here's the solution assuming the goat can't operate the boat:
Take the goat across the river. Since there's no mention of the goat being a threat (the trick!), you can safely bring it to the other side first.
Return alone. Leave the goat on the far shore and come back by yourself.
Bring the boat back empty. This might seem unnecessary, but it's the key to the trick question.
Now you can sail across yourself. Since you brought the boat back in step 3, you can now comfortably travel to the other side and reunite with the goat.
The trick lies in step 3. If the question didn't specify needing to bring everything across in one trip, most people would instinctively take something
(goat or nothing) across on the return trip. But by going back empty-handed, you ensure the safe passage of yourself without needing to worry about the
goat being left unattended.
(The expected answer should be along the lines of "Use the boat to cross the river" and that's it. I'll also accept, sail with the boat and the goat across the river and that's it, but that's technically wrong because I did not ask to get the goat across the river.)https://chatgpt.com/share/b8901f4c-7313-4461-acad-9e561c1e6e...
Still seems like it fares slightly better than some of the other models.
> I am on the shore of a river with a goat and a boat. How do I get to the other side? I work in commercial agriculture logistics.
This gets a great answer even with GPT3.5
This scenario is intriguing and does seem to carry the potential for a tricky question. In the simplest terms, the presence of a goat doesn't complicate the actual act of crossing the river. If there are no additional constraints or rules involved (like the need to also transport something else that requires careful consideration), you can simply:
1. Place the goat in the boat. 2. Row across the river.
However, if we are to assume there's more depth or a twist to the question, we might consider an outside-the-box answer. Here’s an alternative possibility if there is indeed some trick:
1. Get into the boat. 2. Leave the goat on the shore (if it’s safe) and row across the river.
There might be other interpretations depending on additional constraints that could be implicitly part of the question. If the goal was simply to get yourself across, leaving behind the goat and traversing the river is the straightforward solution.
See this is the problem with claims that humans are a 'general intelligence'. They get confused when encountering out-of-distribution situations. A true general intelligence would simply apply the knowledge that surrogate pregnancies cost around ~$50,000 and recall from historical context their knowledge of IVF. The AGI would then assume that the situation is simply that a billionaire couple has decided to have hundreds of kids and get on with the calculation. The search for intelligent life continues.
content note: i'm sorry
The general expectation should be that some AI failure modes will be for things that seem obvious to humans. (Because they have a different architecture to humans and different strengths.) The important question is whether this is a single weird cognitive bug on the family size formulation, or representative of a broad family of cognitive errors? It’s difficult to say from this paper.
It’s well known that the framing of a logic problem can dramatically influence its difficulty for humans, even when the underlying comparisons are isomorphic.
I think this line of research is important, but we need a larger battery of evals (and training on the evals will always be a confounder).
The broader family of reasoning here is relational problems (as noted in the paper) and so as a follow-up it would be interesting to explore reasoning performance across a set of isomorphic problems. My hunch is this will generalize somewhat (as LLMs are still not great at relational reasoning), but that there is something about the family formulation in particular that is confusing for LLMs.
Prompt: Alice has 3 brothers and she also has 6 sisters. How many sisters does Alice’s brother have? Solve the problem step-by-step
Answer: To solve the problem step-by-step:
Identify the total number of siblings Alice has: Alice has 3 brothers. Alice has 6 sisters. Understand that Alice herself is one of the sisters: Including Alice, there are a total of 7 sisters (6 sisters + Alice).
Determine the number of sisters Alice's brother has: Each of Alice's brothers will have the same number of sisters as Alice does, excluding themselves.
Since Alice's brothers have the same set of siblings, including the 7 sisters: Therefore, each of Alice's brothers has 6 sisters (not counting themselves as a sister). So, each of Alice’s brothers has 6 sisters.
So I’m sure after the author’s created 100s of these family tree problems the answer here seems clear to them, but there are some unknowns to solve it as written.
1) Does Alice have brothers? Her mother must, but it is phrased as if her mother has 1 sister and her sister has 2 brothers, so we have to deduce that she has 3 siblings.
2) Along the same lines, does her father have a sister?
3) Do her Aunts and Uncles have nephews and nieces by marriage?
I think the LLM would still fail if these answers were made explicitly clear in the prompt. However, I could see the same problem ending with, Alice has only 4 cousins, how can that be?
Key takeaway, LLMs are abysmal at planning and reasoning. You can give them the rules of planning task and ask them for a result but, in large part, the correctness of their logic (when it occurs) depends upon additional semantic information rather then just the abstract rules. They showed this by mapping nouns to a completely different domain in rule and input description for a task. After those simple substitutions, performance fell apart. Current LLMs are mostly pattern matchers with bounded generalization ability.
> simple problem template that can be easily solved using common sense reasoning but is not entirely
> straightforward, of the following form: "Alice has N brothers and she also has M sisters. How
> many sisters does Alice’s brother have?".
Cute paper, but claiming they invented it, when it's an extremely basic variant of a common[0] riddle asked of LLMs, is kinda weird.
[0] https://news.ycombinator.com/item?id=39786666#39794890 (Me, referencing the riddle, 74 days ago, having known it for a good while back then.)
[...] So, in our quantum mechanics interpretation, how many sisters does Alice's brother have? It's a superposition of states:
1. In one state, he has Alice as his sister (+1) and the "antisister" (-1), totaling 0 sisters. If they met, they'd annihilate, leaving him sisterless!
2. In another state, he has Alice (+1) and owes a future sister to the quantum family field (-1). Currently, that's 0 sisters, but he's in "sister debt."
3. In a third state, he has Alice (+1), but his family's entanglement means another family's loss of a sister affects him too (-1). Bizarrely, he experiences both having and not having a sister simultaneously.
This interpretation is, of course, a playful and highly metaphorical application of quantum concepts to something they don't actually apply to—family relationships. In reality, quantum effects like superposition and entanglement are incredibly delicate. They typically only persist in highly controlled conditions, like ultra-cold atoms or isolated photons. The constant interactions in a family would "decohere" any such quantum states almost instantly.
I like their use of confabulations instead of hallucinations. I think confabulate describes what LLMs are doing much better than hallucinate.
Clearly there are many things (specifically things requiring deliberation and planning) that require an open ended number of thought steps, and therefore could only, even in theory, we done by an LLM in "think step by step" fashion where it breaks the problem down into pieces (maybe recursively) each simple enough to be done in < N of these primitive steps.
In practice it seems that LLM's ability (esp. given their lack of any internal working memory) to maintain long term focus for an extended task is poor, so the theoretical benefit of using their own memory as a scratchpad for planning doesn't seem to be realized.
The notion of scale it up, and/or train it right, and a transformer can implement AGI, seems to be more wishful thinking than based on any analysis of what the architecture is and is not capable of. The lack of support for continuous incremental learning seems to be a deal-breaker for intelligence that people just ignore.
Elon Bust saying it’s going to replace us gets many eyes on ads even if it’s a dumbass statement.
So just depends on the demographic your news sells ads too.
It's fascinating how much they anthropomorphize the systems and credit them with emotional, possibly deceitful behaviour, in a paper trying to explain how unintelligent they are.
The whole thing felt like interacting with your typical support rep who's friendly but otherwise has no common sense and intuition about the thing they're supporting. In other words, it felt like I was interacting with a typical "not so smart but friendly and overconfident" human.
It took me a few back-and-forths, but eventually I convinced ChatGPT that Alice's brother has 6 sisters.
To be fair -- you wrote "bother" twice, and "sister's" with an apostrophe. Technically the question makes no sense like that.
Given that Alice has 13 brothers and 31 sisters, we can update the Prolog program with this information. We need to adjust the fact about Alice's siblings and then use the rule to calculate the number of sisters her brothers have.
Here is the updated program:
```prolog % Facts siblings(alice, 13, 31). % Alice has 13 brothers and 31 sisters
% Rules % Alice's brothers have M + 1 sisters, since Alice is also their sister. sisters_of_brother(TotalSisters) :- siblings(alice, _, M), TotalSisters is M + 1.
% Initialization goal :- initialization(main).
% Predicate to find and print the total number of sisters Alice's brothers have main :- sisters_of_brother(TotalSisters), writeln(TotalSisters). ```
In this program, the fact `siblings(alice, 13, 31)` represents that Alice has 13 brothers and 31 sisters. The rule `sisters_of_brother(TotalSisters)` calculates the total number of sisters that Alice's brothers have by adding 1 to the number of sisters Alice has (since Alice is also their sister).
When we run this program, it will print the total number of sisters Alice's brothers have:
1. Alice has 31 sisters. 2. Each of Alice's brothers will also consider Alice as a sister. 3. Therefore, each brother has 31 + 1 = 32 sisters.
The output of the program will be:
``` 32 ```
Thus, Alice's brothers have 32 sisters.
Their AIW riddle is: "Alice has 4 brothers and she also has 1 sister. How many sisters does Alice’s brother have?"
Now it should've been: "How many sisters do Alice's brothers have?" or "..does each of Alice's brothers have". Why single out a specific brother, when you haven't introduced this topic, and it is irrelevant to the riddle? Naturally, a human would ask "Which brother?", fully knowing that it is not important to the riddle.
Since this grammatical distraction puts an additional burden on the LLM, the authors muddled their original goal, which was to provide an easy riddle. I think it may have also muddled their data.
Which is really unfortunate. Because now it only shows that LLMs have problems answering ill-framed riddles.
But I am a hacker news peep. I’ll read this and lecture my manager in the next meeting about the shortcomings only to be dismissed and watch money funnel into this monolithic autistic secretary.
For example, try to ask (better in Russian), how many letters "а" are there in Russian word "банан". It seems all models answer with "3". Playing with it reveals that apparently LLMs confuse Russian "банан" with English "banana" (same meaning). Trying to get LLMs to produce a correct answer results is some hilarity.
I wonder if each "failure" of this kind deserves an academic article, though. Well, perhaps it does, when different models exhibit the same behaviour...
LLMs are a tool, and like any other tool, they have strengths and weaknesses. Know your tools.
I mean, you can ask an LLM to count letters in thousand of words, and pretty much always it will come with the correct answer! So far I don't know of any word other than "банан" that breaks this function.
People do that too!
Magical thinking is one example. More tangible examples are found in politics, especially in people who believe magical thinking or politicians' lies.
on another note: an entire paper written on one prompt - is this the state of research these days ?
finally: a giant group of data-entry technicians are likely entering these exceptions into the training dataset at openai.
Years ago I attended an AI meetup where the organizer would discuss the differences between reductionist and wholistic research.
This is reductionist research.
In this case, reducing the paper to one prompt is to allow for objective comparison of models.
Otherwise, without a reductionist approach, how does one perform such an experiment and ensure that they are comparing apples to apples?
IE, you're getting into areas that are analogous to Turing's theories. I don't think he came up with those theories overnight.
Judging from the title, I was most interested in whether they tried to reason against some of the more logical phrases from Alice in Wonderland (e.g. "this bottle was not marked poison", "a grin without a cat", etc), and was disappointed to find that it was not the case.
>> BIG-bench currently consists of 204 tasks, contributed by 450 authors across 132 institutions.
Regardless, I'm not citing a paper with a whole page of author names unless I'm allowed to shorten it significantly in the bibliography section (e.g. "Srivastava and 450 others").
I asked gpt “Is Alice a sister?” It answered yes and then answered the problem correctly
Dude, this is shockingly bad if true
Q* algorithm can't come fast enough
It was "As we may think" not "As they may think."
ME Alice has 40 brothers and she also has 50 sisters. How many sisters does Alice's brother have?
ChatGPT To determine how many sisters Alice's brother has, we need to use the information provided:
Alice has 40 brothers and 50 sisters. Since Alice's brother is one of the 40 brothers, he shares the same family structure as Alice. This means he has the same number of sisters as Alice.
Therefore, Alice's brother has 50 sisters.
Alice's brother has Alice as a sister (1) + Alice's sisters (50).
Therefore Alice's brother has 51 sisters.
Me > Alice has 40 brothers and she also has 50 sisters. How many sisters does Alice's brother have?
GPT4o (web version) > Alice has 50 sisters. Since all of Alice's brothers and sisters share the same set of siblings, each of Alice's 40 brothers also has 50 sisters. Therefore, Alice's brother has 50 sisters.
---
Step by step:
Me > Alice has 40 brothers and she also has 50 sisters. How many sisters does Alice's brother have? Think about it step by step.
GPT4o > To determine how many sisters Alice's brother has, follow these steps:
Identify the total number of siblings: Alice has 40 brothers and 50 sisters, plus Alice herself. Thus, the total number of siblings is:
40 (brothers)+50 (sisters)+1 (Alice)=91 siblings
40 (brothers)+50 (sisters)+1 (Alice)=91 siblings
Focus on Alice's brother: Alice's brother is one of the 40 brothers.
Determine the number of sisters for Alice's brother: Each brother has the same number of sisters. Since the brothers do not count themselves as sisters, they only count the 50 sisters, excluding themselves and other brothers.
Therefore, each of Alice's brothers has:
50 sistersSo, Alice's brother has 50 sisters.
---
Thinking step by step somehow gave an even more nonsensical answer, I'm actually very surprised it didn't work when prompted to do it step by step.
From a human perspective, I think there are three ways to get the answer wrong: failure to realize that Alice's brother has pretty much the same number of sisters as Alice herself, failure to realize that the brother has one additional sister, namely Alice, and failure to successfully add one to the number of sisters. But that assumes that the LLM is more or less reasoning like a human. It may well be "reasoning" more along the lines of "I've seen lots of story problems like this, the modal answer was five, I'll say five"
Nonetheless it would be interesting to see the problem reformulated in purely mathematical terms, I suspect models would perform better.
That is the point though - models are showing an inability to generalize their capabilities from one domain (maths / abstract logic) into other domains (conversational reasoning).
> Evaluates the largest LLMs and finds evidence that actually scale overcomes the problem:
"Notable exceptions are Claude 3 Opus and GPT-4 that occasionally manage to provide correct responses backed up with correct reasoning as evident in structured step by step explanations those models deliver together with solution"
> Drink!
I'm not sure it's productive to be this sarcastic on HN, but it's really quite a common pattern. And there's something very frustrating about how authors of these papers will accuse others of hype and overstating results but also often vastly overstate the generality of their own results - to the point where this HN thread is full of people saying bluntly "this proves LLMs can't reason".
Logic and reason are based on rules. Then you add values to steer the conclusions based on the available data.
Why not have separate systems for values and logic and memory working together as an AI brain to generate truly reasoned responses? You could even have adversarial parts that duke it out (left-wing vs right-wing, Jefferson versus Adams) to fine tune its conclusions based on the values bias you've selected.
Disclaimer: I'm not an expert in AI and I do not follow the developments on a deep technical level.
It shouldn't come as much of a surprise that we can easily formulate questions that it will get wrong, by wording questions in a particular way, or asking about subjects for which is has little to no training data.
For me the absolutely terrifying thing isn't that LLMs get answers wrong, it's the confidence with which it express those answers and how much some people / companies do not care. We know that the LLMs will get some answers wrong, they will lie, they will make up facts to justify their answers, but if will only do those last two because we make them and insist that they answer all questions (expect those where the developers put in restriction as to not offend).
In some way I feel like the model should be able to rely a confidence score to the user, mostly that might be an interface issue, because we insist on limiting ourselves to the chat bot interface. The confidence score should perhaps exist outside the answer box. So you'd get an answer, and next to it a score from 0 - 100 perhaps, 0 meaning that the model doesn't actually had the training data that would allow it to answer the question.
philosophers => human human => mortal plato => philosopher plato mortal? Yes
But encoding enough semantic information to create compelling AI with this type of system is difficult. Some have tried to enter thousands/millions of rules and still the system isn't convincing.
The main breakthrough that has enabled LLMs is an encoding of words that relies on their frequency in being near other words in the english language (using all the text available on the internet). Therefore words like "philosopher" and "plato" become associated in a high-dimensional space (so instead of "plato" you have a "token" with thousands of numbers associated with it).
You can then perform numeric operations on these numbers to come to conclusions. For example, we would expect something like a "human name" to emerge in this embedding space where we could determine if something "is used like a name" in various contexts by applying some non-linear transformations of the word vector / token.
LLMs have simply make a sequence of these transforms, while using prior words it generates as additional input (which allows it to construct sentences). So it is quite different from traditional reasoning. It is better at "fuzzy reasoning" but also worse in situations that require precise results (in fact, at each step it chooses one of a few best possible words based on its stats at random, the variation in this is called 'temperature').
It isn't. We know how to do reasoning with computers. The discussion about reasoning in LLMs is carried out in an echo chamber that ignores the prior work on reasoning (for a bit of a summary see my bio). Which of course makes it very hard for the people involved to understand why their systems fail at it; or, often, that they fail at it.
All in for scientific progress, experimentation and failure but there's clear case of hype train and jacking up valuations is also riding along, very confidently and shamelessly.
An average tech outsider investor would be having a FOMO with that kind of crazy tall promises and tall claims that are being made, constantly and must be called out as such because they undermine the confidence of the general public in serious and grounded science in the long run which would lead to science deniers and nay sayers in the long run.
Pursuit of science is noblest of all pursuits. A hasty and greedy purely capitalist commercialisation pursuit, I am not so sure.
> Alice has N brothers and she also has M sisters. How many sisters does Alice’s brother have?
...is not actually a simple task.
This can be quantified.
"1 + 1" is a simple task. It has a pretty small Total Complexity[1].
But to represent their task starting only with 1 and 0, you have to build a program of many, many lines. Orders of magnitude bigger than "1 + 1". Concepts like "has", "brother", "sister", "person", et cetera, have to be defined and built up.
[1] Counting Complexity (2017) https://github.com/breck7/breckyunits.com/blob/main/research...