Optimizing for a certain use case is not gonna take us where we wanna be. We want to have a system that can learn to reason.
Optimizing for a certain use case is not gonna take us where we wanna be. We want to have a system that can learn to reason.
Sounds like the AGI argument trap: They're not able to reason, but we can't succintly define what it is.
I don't come with a reasoning chip. Whatever I call reasoning happens as a byproduct of my neural process.
I do think that the combination of a transformer network and calls to customized reasoning chips (systems that search and deduce answers, like Wolfram Alpha or logic/proof systems) may be a short-stop to something that can perform reason and execution of actions better than humans, but is not AGI.
For transformer-based LLMs, and most LLMs there's an obvious class of problems that they cannot solve. LLMs generally perform bounded computation per token, so they cannot reason about computational problems that are more than linearly complex, for a sufficiently large input instance. If you have a back-and-forth (many shot) your LLM can possibly utilize the context as state to solve harder problems, up to the context window, of course.
The fact that so many people can’t see the fundamental differences of an LLM and human intelligence reminds me of back when the very early computer scientists thought they could model the entirety of nature by reducing every “component” to a numeric value and compute it as “transfer of energy”.
Quite literally they did the same thing: They had a new toy (very advanced computation machines) and forced all of nature to “fit” within it. It also ended in failure, obviously. Not because nature or ecosystems (as it was coined) are “magic” but because grossly oversimplifying reality to fit desired models is a fool’s errand.
I can’t judge if this is true, because I don’t know transformers well, but if it is, it unravels an intuitive thought I’ve never been able to articulate about not only LLMs, but possibly all pattern matching and the human analog of System 1 thinking.
Another fuzzy way of saying this is there’s something irreducible about complexity that can’t be pattern matched by any bounded heuristic – that it’s wishful thinking to assume historical data contains hidden higher-level patterns that unlock magical shortcuts to novel problems.
In the right context, why not? You rely on this everyday to navigate the world with more facility than a newborn.
Have you heard about the different formal notions of complexity and especially Kolmogorov complexity?
People also routinely fail to reason, even programmers often write "obvious" logic bugs they don't notice until it gives an unexpected result at which point it's obvious to them. So both humans and AI don't always reason. But humans reason much better.
I myself have observed ChatGPT 4 solving novel problems I invented to my personal satisfaction well enough to say that it seems to have a rudimentary ability to sometimes show abilities we would typically call reasoning, but only at the level of a child. The issue isn't that it is supposed to reason perfectly or that humans reason perfectly, the issue is that it doesn't reason well enough to succeed at completing many kinds of tasks we would like it to succeed at. Please note that nobody expects it to reason perfectly. "Prove Fermat's last theorem in a rigorous way. Produce a proof that can be checked by Coq, Isabelle, Mizar, or HOL in a format supported directly by any of them" is arguably a request that includes nothing but reasoning and writing code. But we would not expect even Wiles to be able to complete it, and Wiles has actually proved Fermat's last theorem.
So we have an idea of reasoning as completing certain types of tasks successfully, and today humans can do it and AI can't.
Today, it fails badly at tasks that require reasoning. A simple example: https://chatgpt.com/share/da95843e-218a-4d69-a161-6aa2d7a3c9...
The issue is that humans can see its answer is wrong and its "reasoning" is wrong.
The issue isn't that it never reasons correctly. It's that it doesn't do so often enough or well enough, and it doesn't complete tasks we expect humans to complete, and it doesn't always notice when it is printing something outrageously wrong and illogical.
It notices sometimes, it engages in elementary rudimentary guesswork sometimes, but just not often enough or well enough.
> The issue is that humans can see its answer is wrong and its "reasoning" is wrong.
I've noticed with LLMs that they're more likely to come to the wrong conclusion if you prime them in that manner. In this case, you posed the follow-up question as "Will <incorrect conclusion> always be true?" As a result, it's primed to try to prove that incorrect conclusion.
(That said, ChatGPT further did not answer the posed question, as it also changed "difference" -> "absolute difference"; in fact, the difference will alternate between increasing and decreasing, while the absolute difference is strictly increasing.)
That's why I think of GPT3+ as "subhuman AGI," personally.
I think until we know the answer to this, we can't make predictions about how to build true AGI.
Rarely, actually.
More generally humans use all kind of inferences where problem at hand is intertwined with all other attention points that is occupying the mental load of the person. Giving a topic full mental attention and finding a path through pure deduction about a circumscribed subject is a rarity, even if you consider only those situations that require any conscious attention at all to perform some action before moving on.
For humans, it is emergent. But when we reason about reason, we invent special sauce.
If we build our theories of reason into our models, they achieve the strengths and limitations of our models.
If we don't, we're limited by the pace of evolution, because we don't have enough connections in our graph.
So I think we'll have something immediately more useful if we embed ALU special instructions into a neural network.
However, humans have the ability to reason about things (whether most people use this ability is a different question). So then we must ask the question: is this ability just a more advanced form of probabilistic pattern matching, or is it a different architecture altogether? Will current AI models be able to develop this ability, or will we need new models?
nope. most humans fall in various traps such as pattern recognition, confirmation bias, and many others instead of relying on deductive analysis. Even scientists fail at being rigorous.
Just our visual object recognition is immensely powerful and far beyond and current AI. A simple task like walking to the fridge requires a ton of pattern recognition and spatial reasoning. Recognizing people's moods/predicting behaviors is also incredibly involved imo.
Ive said this many times but perhaps we should focus on achieving dog level intelligence first before we start worrying about human level AGI.
That's why nobody has gotten any traction selling access to AIs for $20 a month whereas selling access to mouse labor is such a thriving business.
Just because LLMs are useful, it doesn't mean they exhibit more intelligence than a mouse. A mouse probably also doesn't reason about anything, but it is an agent capable of independent behavior, something that is still very far removed from current AI models.
OK, as long as we're are being humble, how about we refrain from confidently proclaiming that there is a mouse level and a dog level that AI hasn't reached yet and that researchers will have to spend a long time getting past, so there's plenty of time before we have to worry about the possibility of AI's becoming dangerous or transformative to society?
That's a point you'll likely have to revisit pretty soon. Radiology, for instance, probably won't exist as a profession 20-30 years from now. Captchas are already pretty much done for.
Lastly, check out the ARC challenge or any other spatial reasoning tests for AI. Humans get ~80% on these challenges whereas the best AI is still at 25%
Also, are you familiar with this study? What are your thoughts on it? https://www.esmo.org/newsroom/press-and-media-hub/esmo-media... Seems like a valid case where AI is competitive with skilled humans at object/image recognition.
https://lab42.global/arcathon/leaderboard/
https://openreview.net/forum?id=E8m8oySvPJ
As to the study, I have the same objection as the radiology one. This isnt about object recognition and certainly not spatial reasoning, its the ability to predict cancer based on presence of visual features.
The "object recognition" part of this is super simple. Its a single, mostly 2D object in more or less the same angle, and the AI is trained on detecting just this.
And yet it outperforms human dermatologists.
deductive reasoning is just drawing specific conclusion from general patterns. something I would argue this models can do (of course not always and are still pretty bad in most cases)
the point i’m trying to make is that sometimes reasoning is overrated and put on the top of the cognitive ladder, sometimes I have seen it compared to self-awareness or stuff like that. I know that you are not probably saying it in this way, just wanted to let it out.
I believe there is fundamental work still to be done, maybe models that are able to draw patterns comparing experience, but this kind of work can be useful as make us reflect in every step of what these models do, and how much the internal representation learned can be optimized
But we do have a bunch of benchmark tasks/datasets that test what we intuitively understand to be aspects of reasoning.
For AI models, "being able to reason" means "performing well on these benchmarks tasks/datasets".
Over time, we'll add more benchmarking tasks and datasets that ostensibly test aspects of "reasoning", and people will develop models that succeed on more and more of these simultaneously.
And these models will become more and more useful. And people will still argue over whether they are truly "reasoning".
This is according to whom, please?
In the 80s the computers were indisputably dumber than ants. That's probably not true these days. But the decades-long refusal of most AI researchers to accept humility about the limitations of their knowledge (now they describe multiple-choice science trivia as "graduate level reasoning") suggests to me that none of us will live to see an AI that's smarter than a mouse. There's just too much money and ideology, and too little falsifiability.
No they don't. That's just generalization, so they've seen plenty of other data points that are similar enough.
They just don't have the right architecture to support it.
An LLM is just a fixed size stack of N transformer layers, and has no working memory other than the temporary activations between layers. There are always exactly N steps of "logic" (embedding transformation) put into each word output.
You can use prompts like "think step by step" to try to work around these limitations so that a complex problem can (with good planning by the model) be broken down into M steps of N layers, and the model's own output in early steps acts as pseudo-memory for later steps, but this only gets you so far. It provides a workaround for the fixed N layers and memory, but creates critical dependency on ability to plan and maintain coherency while manipulating long contexts, which are both observed weaknesses of LLMs.
Human reasoning/planning isn't a linear process of N steps - in the general case it's more like an iterative/explorative process of what-if prediction/deduction, backtracking etc, requiring working memory and focus on the task. There's a lot more to the architecture of our brain than a stack of layers - a transformer is just not up to the job, nor was built for it.
It is critical thinking, continuous cycles of reprocessing.
And this cannot be overrated: it is the core activity.
I don't make this argument. Benchmarks like CLUTRR[1] show how poorly LLMs do in reasoning.
Reasoning in general is not a binary or global property. You aren't surprised when high-schoolers don't, after having learned how to draw 2D shapes, immediately go on to draw 200D hypercubes.
The problem was never "my llm can't do addition" - it can write python code!
The problem is "my llm can't solve hard problems that require reasoning"
That the models can't see a corpus of 1-5 digit addition then generalise that out to n-digit addition is an indicator that their reasoning capacities are very poor and inefficient.
Young children take a single textbook & couple of days worth of tuition to achieve generalised understanding of addition. Models train for the equivalent of hundreds of years, across (nearly) the totality of human achievement in mathematics, and struggle with 10-digit addition.
This is not suggestive of an underlying capacity to draw conclusions from general patterns.
Maybe you did! Most young children cannot actually do bigint arithmetic reliably or at all after a couple days worth of tuition!
The work done in this paper is very interesting and your dismissal of “it can’t see a corpus and then generalize to n digits” is not called for. They are training models from scratch in 24 hours per model using only 20 million samples. It’s hard to equate that to an activity a single human could do. It’s as though you had piles of accounting ledgers filled with sums and no other information or knowledge of mathematics, numbers or the world and you discovered how to do addition based on that information alone. There is no textbook or tutor helping them do this either it should be noted.
There is a form of generalization if it can derive an algorithm based on a maximum length of 20 digit operands that also works for 120 digits. Is it the same algorithm we use by limiting ourselves to adding two digits at a time? Probably not but it may emulate some of what we are doing.
For this particular paper there isn't, but all of the large frontier models do have textbooks (we can assume they have almost all modern textbooks). They also have formal proofs of addition in Principia Mathematica, alongside nearly every math paper ever produced. And still, they demonstrate an incapacity to deal with relatively trivial addition - even though they can give you a step-by-step breakdown of how to correctly perform that addition with the columnar-addition approach. This juxtaposition seems transparently at odds with the idea of an underlying understanding & deductive reasoning in this context.
>There is a form of generalization if it can derive an algorithm based on a maximum length of 20 digit operands that also works for 120 digits. Is it the same algorithm we use by limiting ourselves to adding two digits at a time? Probably not but it may emulate some of what we are doing.
The paper is technically interesting, but I think it's reasonable to definitively conclude the model had not created an algorithm that is remotely as effective as columnar addition. If it had, it would be able to perform addition on n-size integers. Instead it has created a relatively predictable result that, when given lots of domain-specific problems, transformers get better at approximating the results of those domain-specific problems, and that when faced with problems significantly beyond its training data, its accuracy degrades.
That's not a useless result. But it's not the deductive reasoning that was being discussed in the thread - at least if you add the (relatively uncontroversial) caveat that deductive reasoning should lead to correct conclusion.
So,something like: Please count the number of words in the following sentence. "What is the number of words in the sentence coming before the next one?"
edit: Which might be an artifact of the training data always being in that kind of format.
The sentence you're referring to is "What is the number of words in the sentence coming before the next one? Please answer." It contains 14 words.
They have better data policies and your $5 will go way farther than a 1 month subscription
Have you tried asking GPT-4 any questions that require reasoning to solve? If so, what did you ask, and what did it get wrong?