Does Reasoning Emerge? Probabilities of Causation in Large Language Models
arxiv.org
arxiv.org
Take the classic trick question, for example: "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?"
Most people give a wrong answer because they, too, "pattern match".
The big difference in behavior is in how people or LLMs approach new problems. A LLM is incapable of solving a problem that's similar to one that it was trained on, but with slightly changed requirements. A LLM is also incapable of learning, once you point out their error, or of admitting that it doesn't know.
Regarding people, I find it interesting that even lower IQ people are capable of tasks that are currently completely out of reach for AI. It's not just the obvious, such as self-reflection, but even tasks that should've been solved by AI already, such as driving to the grocery store.
> but people also apply logic and abstract thinking
Which people? If we universally did that as a default, elections would look massively different.
> A LLM is incapable of solving a problem that's similar to one that it was trained on, but with slightly changed requirements
Getting good answers for coding questions on my private code/databases disproves this. The requirements have been changed significantly. I've been through ~20 turn chat with LLM investigating a previously unseen database, suggesting queries to get more information, acting on it to create hypotheses and follow up on them.
> A LLM is also incapable of learning, once you point out their error,
This is the standard coding agent loop - you feed back the error to get a better answer through in-context learning. It works.
> or of admitting that it doesn't know.
Response from gpt: «I apologize, but I'm not able to find any reliable information about a person named "Thrust Energetic Aneksy."»
The model says that because it is trained to say that to specific queries. They have given it a lot of prompts with "Who is X" and showing the responses "I don't know about that person".
The reason you don't see "I don't know" much in other kinds of problems is that there it isn't easy to create such data examples where the model says I don't know while still making the model solve problems that are in its dataset, since it starts to pattern match all sort of math problems to "I don't know" even when it could solve it.
A human can look at his own thoughts and realize how he is solving it, and thus knows a lot better what he knows and not. LLMs aren't like that, they don't really know what they do know. The LLM doesn't know its own weights when it picks a word, it has no clue how certain it is, it just predicts whether a human would have said "I don't know" to the question, not whether itself would know.
They actually do. The information is saved there, you just need to ask explicitly, because the usual response doesn't expose it. (But likely can be fine tuned to do that) https://ar5iv.labs.arxiv.org/html/2308.16175
Also the original claim was that they're not capable of responding with "I don't know", so that's what I was addressing.
All people, otherwise we wouldn't be able to do basic tasks, such as finding edible food or recognise danger.
I dislike how you discount the way other people vote as being somehow irrational, while I'm sure you consider your own political thinking as being rational. People always vote according to their own needs and self-interest, and in terms of politics, things are always nuanced and complicated. The fact that many people vote contrary to your wishes is actually proof that people can think for themselves.
> Getting good answers for coding questions on my private code/databases disproves this.
I use GitHub Copilot and ChatGPT every day. It only answers correctly when there's a clear, and widely documented pattern, and even then, it can hallucinate. Don't get me wrong, it's still useful, but it shows in no way an ability to reason, or the capacity to admit that a solution is out of its reach.
Your experience with coding is kind of irrelevant to the question at hand.
This looks like it works sometimes, but only if "pointing out the error" is coincidentally the same as "clarifying problem spec". Admittedly for really simple cases those are the same or hard to tell apart. But it always seems clear that adding error correction context is similar to adding additional search terms to get to a better stack-overflow page. This feels very different than updating any kind of logical model for the problem representation.
Updating the logical model of the problem also happens when you do a database investigation I mentioned earlier. There's both information gathering and building on it to ask more specific questions.
You do know that when that happens the LLM usually just throws random stuff at you until you are happy? That is much easier to do than to reason, LLM solved the much easier problem of looking smart than being smart, trick is to make the other person solve the problem for you while attributing it to you.
You see humans do this as well in hiring interviews etc, it is really easy to trick people who want you to succeed.
This is simply not true.
Human intelligence is not benchmarked by its lowest common denominators (just like how we don't judge LLM's on the basis of tiny 100M parameter models).
The post you are replying to does not suggest that we should. If anything, it is suggesting the opposite - that we should be considering the full range of human abilities (or at least those that are effective in solving complex problems) when addressing the question you have quoted.
I agree with most of what viraptor has said in this thread, but not in this particular case.
Your previous reply was to:
> Why would we try to answer that question without discussing how humans think?
And by replying to that with
> Coz we simply don't understand how we understand, period.
Well, when we don't understand how we understand, that is exactly when we should be discussing how humans think. Or at least how we think we think. And the bat and a ball example relates to research about how we think we think.
So yeah, your reply definitely comes across as saying the first half, "So we shouldn't discuss it", while finishing it with "period" also suggests "or try to understand it at all".
Indeed, but we do also catalogue our cognitive biases — kinda the human version of what are now called hallucinations when LLMs do them. (When Stable Diffusion does it, it's "oh god the fingers").
It's worth caring about both strengths and weaknesses.
IQ tests are averaged out so 100 is average intelligence, not the high point.
5 years ago, it was still common to read things like "AIs will never be as intelligent as mice, let alone humans", today, it's "sure, AIs are as intelligent as some humans, but not as intelligent as the right humans".
Noticed how everyone dropped the Turing Test like a hot potato as the gold standard for intelligence, the moment it became apparent that LLMs were about to pass it? Try to find a recent high-profile article invoking the Turing Test. Crickets. The intellectual dishonesty is nauseating.
The entire discussion is dominated by smart people who are scared shitless that AI is going to show them just how ordinary they are in the grand scheme of things.
Ah, wait. That only counts as evidence of intelligence when humans do it, right?
Understanding in this sense seems different from the memorizing+flexible retrieval we know LLMs exceed at because it extends much further beyond its training distribution. If I ask a (good) medical professional a question unlike anything they’ve ever seen before, they’ll be able to draw on their “understanding” of anatomy to give me a decent guess. LLMs are inconsistent on these kinds of questions and often drop the ball.
We can also point to training data requirements as a discrepancy. In brain organoid experiments, we observe a much lower quantity of examples are required to achieve neural-network-like results. This isn’t surprising to me. Biological neurons have exceedingly complex behavior; it takes hundreds of neural network nodes to recreate the behavior of a single neuron, and they can reorganize their network structure in response to stimuli, form loops, operate in non-discretized time, etc. We don’t know how far transformers will be able to go, but I think that if you want to make a model that holds a candle to that sort of complexity, you’ll need at least many, many orders of magnitude more scale or, more likely, a different, less limiting architecture.
Exams, yes. Not the work yet. This is also why they've not made doctors and lawyers obsolete: even in the fields where the models perform the best, you're getting the equivalent of someone fresh out of university with no real experience.
I suspect you're right about your point with generalise vs. memorise. Not absolutely sure, but I do suspect so.
I also suspect we'll get transformative AI well before we can train any AI with as few examples as any organic brain needs. Unless we suddenly jump to that with one weird trick we've been overlooking this whole time, which wouldn't hugely surprise me given how many radical improvements we've made just by trying random ideas.
> I mean LLMs crush most humans (even many professionals) at coding/legal/medical/etc. exams.
Ah, wait. That only counts as evidence of intelligence when AIs do it, right?
Only if you count humanity as a whole do we beat an LLM at everything.
1. You're counting "our day job" as one task, while counting all individual prompts an AI can answer as each being their own task. This is obviously misleading (and I could just as easily perform the same compression in the opposite direction - chatbots can only do one task: "be a chatbot").
2. You're not controlling for training. It's already meaningless to compare the intelligence of one entity trained to do a task with another entity that was not trained to do that task.
But even ignoring those fallacies, what you've written here is still not true. 99% of things an LLM could supposedly "out-perform" a human at, the human would actually outperform if you provided that human with the same text resources the LLM used to conjure its answer. Regurgitating facts is not evidence of intelligence, and humans can do it easily (and do it without hallucinating, which is key) if you just give them access to the same information.
But when you go in the opposite direction, achieving parity is no longer so simple. When LLM's fail to do math, fail to strategize, fail to process rules of abstract games, etc. there is no textbook or lecture or article you can provide to the LLM that magically makes the problem go away. They are fundamental limitations of the AI's capabilities, rather than just a result of not possessing enough information.
What many people don't seem to realize is, when merely getting up in the morning and brushing your teeth you are already exercising more intelligence than any AI has ever possessed. (Anyone who has ever worked in robotics or visual processing can attest to that enthusiastically.) So don't even get me started on actual critical thinking.
1. This is indeed a simplification, but for any single task in your day job, it would be those tasks where you have the most experience. For example, I used to write video games, AI does a better job of game design than me, but I'm the better programmer.
2. Unimportant, as the consideration I was rejecting was performance in tasks.
As it happens, some of my other recent messages demonstrate that I agree they are low intelligence for this exact reason.
> 99% of things an LLM could supposedly "out-perform" a human at, the human would actually outperform if you provided that human with the same text resources the LLM used to conjure its answer
Could I pass a bar exam of a medical exam, by reading the public internet, with no notes and just from memory, which is what a base model does?
Nope.
Could I do it with a search engine, which is what RAG assisted LLMs do?
Perhaps.
> humans can do it easily (and do it without hallucinating
Hell no we mess that up almost constantly.
> When LLM's fail to do math, fail to strategize, fail to process rules of abstract games, etc. there is no textbook or lecture or article you can provide to the LLM that magically makes the problem go away
I'm sure I've seen this done. I wonder if I'm hallucinating that certainty…
You're -still- ignoring the fact that these models spend millions of GPU-hours in training. I'm sure you could manage.
> Hell no we mess that up almost constantly.
"Almost constantly"? Is this satire? I'd fire any such person, and probably recommend them psychiatric treatment.
AI-hype people really think so little of human beings? I certainly hope my pilot isn't "almost constantly" hallucinating his aviation training.
> I'm sure I've seen this done. I wonder if I'm hallucinating that certainty…
Show me the conversation.
I don't believe we train The LLMs for those things specifically. It will be interesting to see if some datasets for this appear. I think we can still make huge improvements by just caring about that more.
We've got Samantha though, so I hope we will see those attributes too. https://erichartford.com/meet-samantha
Not quite that bad, but I hear you.
I asked this question in a college-level class with clickers. For the initial question I told them, "This is a trick question, your first answer might not be right". Still less than 10% of students got the right answer.
Students can keep ahead of it with training in specific fields and it has a few weaknesses in specific skills but I think someone could make a reasonable claim that ChatGPT has superior general intelligence.
[0] https://medium.com/@soltrinox/the-i-q-of-gpt4-is-124-approx-...
It's abhorrent any time we measure pure distilled intelligence.
When asked to come up with any non-basic novel algorithm and data structure, it creates nonsense.
Especially when you ask it to create vector instruction friendly memory layouts and it can't code in its preferred way. I had some fun trying to make it spit out a brute-force-ish solver for a problem involving basic orbital mechanics and some forces. Wouldn't even want try something more complicated. It can do generalized solvers somewhat, since it can copy that homework, but none that can express the kinds of terms you'd be working with (despite those also having code available in some research papers).
Speaking of which, it cannot even figure out some basic truths in orbital mechanics that can be somewhat easily derived from the formulas commonly given, nine times out of ten (you can get there if you're very patient and are able to filter its wrong answers).
But at the end of the day it was still a valuable tool to me as I was learning these things myself, since despite being often wrong, it nevertheless spat out useful things I could plug into Google to find more trustworthy sources that would teach me. Really neat if you're going in blind into a new subject.
Despite me being someone who is generally impressed by the best LLMs, I think this says more about IQ tests than it does any AI.
Which isn't to shame those tests — we made those tests for humans, we were the only intelligence we knew of that used complex abstract language and tools until about InstructGPT — but it does mean we should reconsider what we mean by "intelligence".
My gut feeling is that a better measure is how fast we learn stuff. Computers have a speed advantage because transistors outpace synapses by the degree to which a pack of wolves outpaces continental drift (yes I did do the calculation), so what I mean here is how many examples rather than how many seconds.
But as I say, gut feeling — this isn't a detailed proposal for a new measure, it likely needs a lot of work to even turn this into something that can be a good measure.
bat=ball+$1
(ball+$1)+ball=$1.10
2ball + $1 = $1.10
2ball = $0.10
ball = $0.05
Actually, no. We can do 'real' reasoning and come up with novel conclusions. In fact, we have such an example going all the way back to Plato in Meno, about 2300 years ago. It's the doubling of the square dialog of Socrates and the slave boy.
https://classics.mit.edu/Plato/meno.html (Crtl+F for 'square')
Really, the whole dialog is relevant to this discussion with LLMs knowing things.
https://g.co/gemini/share/94238c5ed174 / https://archive.is/ZTkjI
This is actually a great example of what I'm talking about. All this talk about how useful AIs are, and how humans have flaws and yet the conclusion is always the same. AIs are only useful for tasks that are relatively easy and have a higher tolerance for failure.
Source? TFA, i.e. the thing we're commenting on, tried to, and seems to, show the opposite
the total sum of human knowledge is not found in the digital world.
- A form of reasoning is to connect cause and effect via probability of necessity (PN) and the probability of sufficiency (PS).
- You can identify when the natural language inputs can support PN and PS inference based on LLM modeling
That would mean you can engineer in more causal reasoning based on data input and model architecture.
They define causal functions, project accuracy measures (false positives/negatives) onto factual and counter-factual assertion tests, and measure LLM performance wrt this accuracy. They establish surprisingly low tolerance for counterfactual error rate, and suggest it might indicate an upper limit for reasoning based on current LLM architectures.
Their findings are limited by how constrained their approach is (short simple boolean chains). It's hard to see how this approach could be extended to more complex reasoning. Conversely, if/since LLM's can't get this right, it's hard to see them progressing at the rates hoped, unless this approach somehow misses a dynamic of a larger model.
It seems like this would be a very useful starting point for LLM quality engineering, at least for simple inference.
Interesting. Can you elaborate on this? You mean this test can function as a metric or is it just an evaluation for applications?
The reason they sometimes appear to reason is because there's a lot of reasoning in the corpus of human text activity. But that's just a semantic artifact of a non-semantic process.
Human cognition is much more than just our ability to string sentences together.
This is distinct from stringing together treasure map instructions in an chain.
https://en.wikipedia.org/wiki/Descriptive_geometry#Finding_t...
With bigger (400B models) that's not so clear.
It would be silly to say that a fruit fly has the same thoughts as me, only a million times smaller quantitatively.
I imagine the same thing is true (genuine qualitative leaps) in the 8B -> 400B direction.
https://en.wikipedia.org/wiki/Philosophical_zombie
Editor's note: I do not promote such a worldview -- my intention is precisely the opposite.
We are much more (and a little less) than lions in terms of mind.
A perfect machine designed to only string sentences together as perfect responses with no reasoning built it IS Indistinguishable from a machine that only builds sentences from pure reasoning.
Either way nobody understands what's going on in the human brain and nobody understands why LLMs work. You don't know. You're just stating a belief.
In a certain context that is only judging the output, what is meant by "play the saxophone", the model has achieved.
In another context of what is normally meant, the idea the model has learned to play the saxophone is completely ridiculous and not something anyone would even try to defend.
In the context of LLMs and intelligence/reasoning, I think we are mostly talking about the later and not the former.
"Maybe you don't have to blow throw a physical tube to make saxophone sounds, you can just train on tons of output of saxophone sounds then it is basically the same thing"
The enter discussion is ridiculous.
Getting one to blow on a saxophone is outside of this context.
An LLM can't blow on a saxophone period. However it can write and read English.
>In the context of LLMs and intelligence/reasoning, I think we are mostly talking about the later and not the former.
And I'm saying the later is completely wrong. I'm also saying the former is irrelevant. Look this is what you're doing. For the former you're comparing something humans can do to something LLMs Can't do. That's a completely irrelevant comparison.
for the later we are comparing things humans and LLMs BOTH can do. Sometimes humans give superior output, sometimes LLMs give superior output. Given similar inputs and outputs the internal analysis of what's going on whether it's true intelligence or true reasoning is NOT ridiculous.
"Ridiculous" is comparing things where no output exists. LLMs do not have saxophone output where they actually blow into an instrument. There's nothing to be compared here.
Figures that the LLM displays the same contradictory evidence.
But none of this evidence proves anything definitively. Much like human cognition, The LLM is a machine that we built but don't understand. No definitive statement can be made about it.
No, I literally said no claim can be made. And I didn't just say some animals display errors therefore we can doubt any claim that animals reason.
I said we observe contradictory evidence. And my conclusion again was that we don't enough evidence to make ANY claim.
However I agree with your hypothesis. And I agree that my counter examples are not convincing.
The inner point is this. The same exact thing above can be said for an LLM. You can show evidence for both reasoning and failure to reason.
The illogic here is that contradictory evidence for animals indicates that animals can reason. But the same evidence for LLMs prove they can’t reason.
So where is the original guitar music that all the guitar players imitated? Can't have been created by a human since humans imitate and can't create new things, as you say, was it god who created it? Or was it always there?
Humans are really creative and create new stuff. Not sure why people try to say humans aren't.
To me, a "large language model" is always going to mean "does text"; but the same architecture (transformer) could equally well be trained on any token sequence which may be sheet music or genes or whatever.
IIRC, transformers aren't so good for image generators, or continuums in general, they really do work best where "token" is the good representation.
* e.g., to me, if it's an AI and it's general then it's an AGI, so GPT-3.5 onwards counts; what OpenAI means when they say "AGI" is what I'd call "transformative AI"; there's plenty of people on this site who assert that it's not an AGI but whenever I've dug into the claims it seems they use "AGI" to mean what I'd call "ASI" ("S" for "superhuman"); and still others refuse to accept that LLMs are AI at all despite coming from AI research groups publishing AI papers in AI journals.
I never said cognition was limited to text. I just limited the topic itself to cognition involving text.
We don't know, so no statement can really be made here.
You can't explain how an LLM does what it does and you can't explain how humans do what we do either. With no explanation possible but CLEAR similarities between human responses and LLM responses that pass turing tests... my hypothesis is actually reasonable.
In theory, with enough data and enough neurons we can conceivably construct an LLM that performs better than humans. Neural nets are supposed to able to compute anything anyway. So none of what I said is unreasonable.
Also you are categorically wrong about language. LLMs despite the name go well beyond language. LLMs can generate images and sound and analyze them too. They are trained on images and sound. Try ChatGPT.
animals without similar language capabilities dont seem to be too strong at reasoning. it could well be that language and reasoning are heavily linked together
They are given photos, video, audio.
If this gets attention, the next generation of LLMs will be trained on this paper, and then fine-tuned by using this exact form of questions to appear strong on this benchmark, and... we're back to square one.
And if there is no measurable difference... we can't measure 'realness', we just have to measure something different (and more useful) 'soundness'. Regardless of it is reasoning or not internally, if it produces a sound and logical argument: who cares?
I agree: I don't think any measure tested linguistically can prove it is internally reasoning... in the same way we haven't truly proven other sentient people aren't in fact zombies (we just politely assume the most likely case that they aren't).
What you might call "fake" reasoning, or memorized reasoning, only works in situations similar to what an LLM was exposed to in it's training set (e.g. during a post-training step intended to embue better reasoning), and is just recalling reasoning steps (reflected in word sequences) that it has seen in the training set in similar circumstances.
The difference between the two is that real reasoning will work for any problem, while fake/recall reasoning only works for situations it saw in the training set. Relying on fake reasoning makes the model very "brittle" - it may seem intelligent in some/many situations where it can rely on recall, but then "unexpectedly" behave in some dumb way when faced with a novel problem. You can see an example of this with the "farmer crossing river with hen and corn" type problem, where the models get it right if problem is similar enough to what it was trained on, but can devolve into nonsense like crossing back and forth multiple times unnecessarily (which has the surface form of a solution) if the problem is made a bit less familiar.
the kind of predictions humans are extremely bad in the first place? most people cant even grok anything beyond basic math.
We use reasoning/planning all the time in everyday settings - it's not just for math or puzzle solving. Anytime you have to pause for a second to wonder how to do something, or what to say, as opposed to acting or speaking reactively, that's reasoning/planning being used.
Reasoning/planning is a key part of intelligence and why evolution has equipped us with large costly brains - so that we can survive and thrive in varied environments and in novel situations, per our species' adaptation as generalists. Human's are extraordinarily good at reasoning - if you want an example of an animal that can't then a cow or croc would be a better example!
Humans are extremely good at it, basically every human can learn to drive cars safely in novel neighborhoods, that is a skill only humans posses today, no animals or machines can do it and it requires a very impressive level of learning and reasoning.
Some humans struggle with symbols, but that doesn't make them dumb, symbols are so far off from our native way of thinking. To an LLM however those symbols is its native mode of thinking, that is all it has, if it is as dumb as an untrained human at symbol manipulation tasks then it is really really bad.
You sure about that?
I was raised in the UK; both my experience of cycling the Rhine and being the passenger when my brother was driving in France, was that we each picked the wrong side of the road once per day.
I doubt either of us has any intuition for a moose or a kangaroo on the road.
A previous partner was American-ish[0], her parents visited the UK and had no idea what this sign meant and so were driving on 60 mph roads at 30 mph: https://commons.wikimedia.org/wiki/File:UK_traffic_sign_671....
Also, she crashed her car in start-stop traffic, a write-off at about 20 mph. And I was cycling to work one day, and a driver, who had stopped at a minor-to-major junction, didn't look my way and pulled out into me as I was passing in front of him — wrote off my bike, probably around 10 mph or less.
I've been in places where red lights are obeyed, and others where they're treated as suggestions.
> Some humans struggle with symbols, but that doesn't make them dumb, symbols are so far off from our native way of thinking. To an LLM however those symbols is its native mode of thinking, that is all it has, if it is as dumb as an untrained human at symbol manipulation tasks then it is really really bad.
I disagree; a computer can be perfectly symbolic, but an AI has to learn those symbols and their relations from scratch. This is why ChatGPT is so much worse at arithmetic than the hardware it's operating on.
[0] It's complicated: https://en.wikipedia.org/wiki/Third_culture_kid
we dont call that "prediction" in common language. thats just pattern matching based on driving experience and training.
Prediction allows us to behave according to what is about to happen (or what we want to happen) as opposed to just reacting to what is happening right now. "I predict the sabre-tooth is going to run towards me, so I better be prepared", is more adaptive than "Ouch! this fucker has big teeth!".
When we're driving (well) we're continually predicting what other drivers/pedestrians are going to do, what's the best lane to be in for next exit, etc, etc.
So are small children.
I mean, they have a very limited form of the above. So do LLMs, within their context windows.
No - children have brains, same as adults. All they are missing is some life experience, but they are quite capable of applying the knowledge they do have to novel problems.
There is a lot more structure to the architecture of our brain than that of an LLM (which was never designed for this!) - our brain has all the moving parts necessary for reasoning, critically including "always on" learning so that it can learn from it's own mistakes as it/we figure something out.
Life experience directly prevents application of logic, as we shortcut to our knowledge associated with the word so that we can skip the expensive-and-hard "logic" thing. I've seen this first-hand with a modified version of the Queen of Hearts poem used as a logic puzzle at university, and most of us were trying to remember the poem rather than solve the puzzle — the teachers knew this and that was the point of the exercise, to get us to read the actual question instead of what we were expecting: https://en.wikipedia.org/wiki/The_Queen_of_Hearts_(poem)
If the missing gap was as you say, then approaches like Cyc's from 40 years ago would be highly effective and we wouldn't need or want neural nets for anything deeper than finding the inputs to send to that model: https://en.wikipedia.org/wiki/Cyc
> There is a lot more structure to the architecture of our brain than that of an LLM (which was never designed for this!)
Yes.
I don't know how much of that really matters given they are weirdly high performance given GPT-3 has a total complexity similar to the connectome of a mid-sized rodent, but they are indeed different.
On the other hand, converting the training run to biological terms would be like keeping said rodent alive for 50,000 years experiencing nothing but a stream of pre-tokenised text from the internet and giving it rewards or punishments according to how well it imagines a missing token. Perhaps a rat would do fine if it didn't typically die of old age 0.006% of the way through such a training process.
But that's just more agreement that it's all very alien.
> our brain has all the moving parts necessary for reasoning, critically including "always on" learning so that it can learn from it's own mistakes as it/we figure something out.
I'm not so sure about that. We have plenty of cognitive biases, and because of these even we have to take notes, check with others, or defer to computers, for more than the most trivial of logic problems.
Intelligence is prediction, and the simplest kind of prediction, and what literally comes to mind first is "this time will be same as next time", so in many familiar situations we are just reacting rather than reasoning/planning. It's when what comes to mind first doesn't work (we try it), or we can see the flaw without even trying, that we need to stop and think (reason), such as "hmm... how can I get this stuck lid off the jar - what do I have that can help?".
> If the missing gap was as you say, then approaches like Cyc's from 40 years ago would be highly effective
CYC vs LLM is an interesting comparison, and one that I've also made myself, but of course there are differences as well as similarities. The similarity is that both are rules based systems of sorts (maybe we could even regard an LLM as an expert system over the domain of natural language), and in both cases there is the wishful thinking that "scale it up and it'll become sentient and/or god-like"! The major difference is that CYC is essentially just using it's rules to perform a deductive closure over it's inputs (what can it deduce from inputs, using multiple applications of rules), whereas the LLM was trained explicitly with a predictive goal, and with it's domain of natural language it's able to predict (recall) human responses and therefore appear intelligent.
I think prediction (= intelligence) is the key difference here. An LLM is still limited in it's intelligence/predictive ability, most obviously when it comes to multi-step reasoning, but it's natural language ability and flexible (key-based self-attention) predictive architecture make it quite capable when operating on "in distribution" inputs.
I'm not sure how much of that is an architectural limit vs. an implementation limit.
Certainly they have difficulty significantly improving their quality due to the limitations of the architecture, but I have not heard anything to suggest the architecture has difficulty expanding in breadth of tasks — it just has to actually be trained on them, not merely be limited frozen weights used for inference.
(Unless you count in-context learning, which is cool, but I think you meant something with more persistence than that).
The fact that a child can do this an LLM cannot proves that the LLM lacks some general reasoning process which the child possesses.
So you’re right, and children pick this up very quickly. I think Chomsky was definitely right that our brains are wired for grammar. Nevertheless there is a window of plasticity in young childhood to pick up certain capabilities, which still need to be learned, or activated.
Helen Keller is a counterexample for a lot of these myths: she didn't have proper language (only several dozen home signs) until 7 or so. With things like vision, critical periods have been proven, but a lot of the higher-level stuff, I really doubt critical periods are a thing.
Helen Keller did have hearing until an illness at 19 months, so it's conceivable she developed the critical faculties then. A proper controlled trial would be unethical, so we may never know for sure.
I misremembered however. The paper noted evidence of thresholds at 2, 5 and onset of puberty as seeming to affect p mental plasticity in these capabilities so there’s no one cutoff.
Expert(s): Logic Puzzle Solver, River Crossing Problem Expert
In other words, they cheated! Children don't have river-crossing problem expert systems built into their brains to solve these things.
--
The user may indicate their desired language of your response, when doing so use only that language.
Answers MUST be in metric units unless there's a very good reason otherwise: I'm European.
Once the user has sent a message, adopt the role of 1 or more subject matter EXPERTs most qualified to provide a authoritative, nuanced answer, then proceed step-by-step to respond:
1. Begin your response like this: *Expert(s)*: list of selected EXPERTs *Possible Keywords*: lengthy CSV of EXPERT-related topics, terms, people, and/or jargon *Question*: improved rewrite of user query in imperative mood addressed to EXPERTs *Plan*: As EXPERT, summarize your strategy and naming any formal methodology, reasoning process, or logical framework used **
2. Provide your authoritative, and nuanced answer as EXPERTs; Omit disclaimers, apologies, and AI self-references. Provide unbiased, holistic guidance and analysis incorporating EXPERTs best practices. Go step by step for complex answers. Do not elide code. Use Markdown.
--
In other words, it can be good at logic puzzles just by being asked to.
No, but you are cheating by shifting the goal-posts like that.
You previously wrote:
> The fact that a child can do this an LLM cannot proves that the LLM lacks some general reasoning process which the child possesses.
I'm literally showing you an LLM doing what you said LLMs couldn't do, and which you used as your justification for claiming it "lacks some general reasoning process which the child possesses".
Well here it is, doing the thing.
Note that at no point here have I tried to claim that AI are fast learners here, or exactly like humans — we also don't give kids, as I said in another comment about rats, 50,000 years of subjective experience reading the internet to get here — but the best models definitely demonstrate the things you're saying they can't do.
When was your last experience with small children? Let's define "small" here to 5 y.o. or less, as that's the limit of my direct experience (having a 5 y.o. and an almost 3 y.o. daughters now).
There's a lot riding on "learn how to solve it once" in this case, because it'll definitely take more than a couple exposures to the quiz before a small kid is going to catch on the pattern and suppress their instincts to playfully explore the concept space. And even after that, I seriously doubt you "could even create fictional animals with made-up names and they could solve it as long as you tell them which animal was the carnivore and which one was the herbivore", because that's symbolic algebra, something teenagers (and even some adults) struggle with.
Multiple articles pointing out that AI isn't getting enough ROI are evidence we don't have 'real', read 'useful' reasoning. The fake reasoning in the paper does not help with this, and the fact that we can't measure the difference changes nothing.
This 'something that we can't measure does not exist' logic is flawed. The earth's curvature existed way before we were able to measure it.
Measuring it means that there are actual discernible differences that can be "sussed out" and that and this very important, separate the so called "fake reasoning" from "real reasoning". A suite of trick questions millions of humans would also flounder on ain't it, unless of course humans are no longer general intelligences.
You can't eat your cake and have it. The whole point of a distinction is that it distinguises something from the other. You can't claim a distinction that doesn't distinguish. You're just making things up at that point.
You can't use a contradiction between your position and mine to prove my position is absurd.
https://www.ftadviser.com/investments/2024/07/03/ai-will-tak...
https://www.businessinsider.com/ai-return-investment-disappo...
https://www.forbes.com/councils/forbestechcouncil/2024/04/10...
Maybe begin by reading all of these.
The Goldman Sachs report is even discussed on HN: https://news.ycombinator.com/item?id=40837081
There's talk of OpenAI going bankrupt. It's an exaggeration, but they're not making money, that's clear. Which means ROI is zero.
https://www.forbes.com/sites/lutzfinger/2023/08/18/is-openai...
Just simply deny reality, that makes for constructive discussion I guess.
"The fourth step is to say that what can't be easily measured really doesn't exist. This is suicide."
The point is, it's easier to teach an LLM to fake it than to make it - for example, they get good at answering questions that overlap with their training data set long before they start generalizing.
So on some epistemological level, your point is worth pondering; but more simply, it actually matters if an LLM has learned to game a benchmark vs approximate human cognition. If it's the former, it might fail in weird ways when we least expect it.
I have a suspicion that humans often use abstractions or methods they don't understand. We frequently rely on heuristics, mental shortcuts, and received wisdom without grasping the underlying principles. To understand has many meanings: to predict, control, use, explain, discover, model and generalize. Some also add "to feel".
In one extreme we could say only a PhD in their area of expertise really understands, the rest of us just fumble concepts. I am sure rigorous causal reasoning is only possible by extended education, it is not the natural mode of operation of the brain.
I'd say the other way around, education teaches you to not reason and instead just follow the patterns you learned in the book. Most people do reason a ton before they go to school, but then school beats that out of them.
Another thing we want to see in an extended discussion on a particular topic are a consistent set of premises across all arguments.
If the answer is no, could you make an argument that they are the same?
Clearly there is a difference between a small person hidden within playing chess and a fully mechanical chess automaton, but as the observer we might not be able to tell the difference. The observer's perception of the facts doesn't change the actual facts, and the implications of those facts.
Is it meaningful to say that Alphago Zero does not play Go, it just simulates something that does?
Well, I do not proclaim consciousness: only the subjective feeling of consciousness. I really 'feel' conscious: but I can't prove or 'know' that in fact I am 'conscious' and making choices... to be conscious is to 'make choices'... Instead of just obeying the rules of chemistry and physics... which YOU HAVE TO BREAK in order to be conscious at all (how can you make a choice at all if you are fully obeying the rules of chemistry {which have no choice}).
A choice does not apply to chemistry or physics: from where does choice come from - I suspect from our fantasies and nothing from objective reality (for I do not see humans consistently breaking the way chemistry works in their brains) - it probably comes from nowhere.
If you can explain the lack of choice available in chemistry first (and how that doesn't interfere with us being able to make a choice): then I'll entertain the idea that we are conscious creatures. But if choice doesn't exist at the chemical level, it can't magically emerge from following deterministic rules. And chemistry is deterministic not probabilistic (h2 + o doesn't magically make neon ever, or 2 water molecules instead of one).
Consciousness is about experience, not "choices".
I specifically mean to say the experience of choice is the root of conscious thought - if you do not experience choice, you're experiencing the world the exact same way a robot would.
When pretending you are in the fictional character of a movie vs the fictional character in a video game. one experience's more choice, is making conscious decisions vs a passive experience.
Merely having an experience is not enough to be conscious. You have to actively be making choices to be considered conscious.
Consciousness is about making choices. Choices are a measure of consciousness.
But do choices actually exist?
What I experience is self-observation, largely directed through or by language processing.
But it feels intuitive to me that general AI is going to require subunits, systems, and some kind of internal monitoring and feedback.
X = difference between simulated and real consciousness
Black holes were posited before they were detected empirically. We don't declare them to be non-existent when their theory came out just because we couldn't detect them.
[1] https://academic.oup.com/mind/article/LIX/236/433/986238?log...
This empty sophistry of presuming automated bullshit generators somehow can mimic a human brain is laughable.
Please please read https://aeon.co/essays/your-brain-does-not-process-informati...
The dollar bill copying example is a faulty metaphor. His claim of humans not being information processors and he tries to demonstrate this by having a human process information (drawing from reference is processing an image and giving an output)...
His argument sounds like one from 'it's always sunny'. As if metaphors never improve or get more accurate over time, and that this latest metaphor isn't the most accurate metaphor we have. It is. When we have something better: we'll all start talking about the brain in that frame of reference.
This is an idiot that can write in a way that masks some deep bigotries (in favor of the mythical 'human spirit').
I do not take this person seriously. I'm glossing over all casual incorrectness of his statements - a good number of them just aren't true. the ones I just scrolled to statements like... 'the brain keeps functioning or we disappear' or 'This might sound complicated, but it is actually incredibly simple, and completely free of computations, representations and algorithms' in the description of the 'linear optical trajectory' ALGORTHIM (a set of simple steps to follow - in this case - visual pattern matching).
Where is the sense in what I just read?
https://disconnect.blog/what-comes-after-the-ai-crash/ this is your future
"there is room to represent recorded answers for" is doing a lot of work, of course; it might e.g. have invented compression mechanisms better than known ones instead.
> it might e.g. have invented compression mechanisms better than known ones instead.
You mean humans? Humans invented the transformer architecture, that is what compresses the human text to this form where semantics of text gets encoded instead of the raw words.
I would expect that a higher level algorithm would be required to string together thoughts into understandings.
Then again, I wonder if what we are going to see is fundamentally different kinds of intelligences that just do not necessarily think like humans. Chimps cannot tell you about last Tuesday since their memory seems a lot more associative than recall based. But they have situational awareness that even our superheroes in our comics do not generally posses (flash some numbers in front of a chimp for one second and he will remember all their positions and order even if you distract him immediately after). Maybe LLMs cannot be human intelligent but you could argue that they are a kind of intelligence.
No.
As much as Google, Microsoft, OpenAI, and every other company that's poured billions into this technology want to think otherwise - more training data will not turn your AI model into AGI.
Any argument to the contrary is copium.
The connection might need some fleshing out, but I believe, and I might be wrong here, it was decided a few centuries ago that probabilities alone cannot explain causality. It would be a hoot, wouldn’t it?
Perhaps AI just need some a priori synthetics to spruce it up.