Statistical or probabilistic reasoning is still reasoning, and indeed essentially mandatory to make sense of the real world, and is what humans do. Indeed the failure of painstakingly hand-trained symbolic reasoning engines to model reality (as opposed to just simple toy worlds) was the entire reason for the rise of neural networks.
The entire point of a neural network is to efficiently compress large amounts of statistical data into a model with as much predictive power as possible. Somewhere within GPT-4's many layers there has to exist something like Bayes networks and other ultra-efficient representations of the vast network of correlations present in the training corpus – otherwise there's no way it could do what it does.
I didn't mean this by "reasoning". To me, Coq does not reason (in the meaning I intended). It checks proofs are well-formed and assists proof writing.
But yes, there's some implementation of logic in Coq. It applies (a specific form of) logic. Again, LLMs output probable text. They don't do logic. They do statistics on language.
Coq works on the actual objects themselves directly, not on the text that presents them.
LLMs usually fail to answer to riddles correctly for this reason. We saw plenty of occurrences of this in recent times. Each time LLMs are benchmarked, people think of trying riddles and it works to some extent because something that looks like these riddles are in the training set but it eventually fails. Humans also make mistakes when solving riddles, but not for the same reasons (the process is not the same, they are not "processing text", the are really processing logic, even if the process is fallible).
You do not know what humans do, no human does, as we don't have such an access to our internal representations. You neither know what LLMs do, no human does, as even though we have full access to their internals, we have little to no idea how to interpret said internals.
The most interesting thing about LLMs is the degree that they have learned to model the real world while only being exposed to words. Because those words are highly correlated not just with each other, but with an external entity unknown to the LLM, it's very likely that by learning to model that external entity, the LLM's word-predicting power will increase. LLMs are likely to have converged on a world-model simply because it makes the most sense given the available evidence.
This is no different from how human scientists create models of things they cannot directly see based on what they can see. And of course, "seeing" is not directly perceiving the external reality either! Sensory data merely correlates with what's really going on, and the brain has learned to model the world – an unknown external entity – based on the correlations that exist in the sensory language.
I'll admit I did this. I tried to frame / convey my perspective on the word "think".
> Not a very convincing argument.
I tried to detail a bit more though.
> You do not know what humans do
No, but I'm pretty sure I don't try to predict the next word when solving a logical problem through statistical text processing. I would probably do something more like this if I were to bullshit someone or give an insipid political speech, but I think I'm quite bad at it.
(nothing to answer on the rest of your comment, it's interesting and I quite agree)
(edit: but see https://news.ycombinator.com/item?id=37912319)
Neither do transformer / deep learning LLMs. They build an internal representation of the text that captures various relationships and structures within the broader context. They are not Markov chains.
Coq parses the text to an abstract representation and so do LLMs.
Compilers translate their abstract interpretation to another programming language, usually native assembly code or bytecode. Coq applies / verifies logical (inductive) inferences on logical objects (rules, proofs, axioms, predicates) (not a Coq expert, details should be checked). LLMs predict next words.
That's an oversimplification. Deep learning LLMs such as ChatGPT are designed to solve natural language processing (NLP) tasks, not just next-word prediction. In order to "predict next words" they create a model and "reason" about it. They are not simple Markov processes. There's a lot happening in that ~100 layer neural network.
It just turns out that good next-word prediction requires understanding the world.
(And you can equally view humans as "next word predictors" anyway.)
> Reason is the capacity of applying logic consciously by drawing conclusions from new or existing information, with the aim of seeking the truth, with more than 10 or more than 3 evidence, 99.9% reliability or 87.5% reliability.
Apart from the words "consciously" (which has 100 or so different definitions, some of which do and some of which do not apply to AI), and "aim" which can be contentious depending upon if you feel that requires consciousness or can include training rewards… I think the best of them is at the lower of those two mysterious percentages I've never seen discussed before.
And that's despite LLMs being wildly bad at, specifically, logic. Though even this specific weakness in general seems like a very easy thing to get around by hybridising it with the computer it has to run on e.g. by telling it how to use a compiler, as logic is the foundation of the arithmetic that the transformers use to do natural language comprehension in the first place.
Edit: yeah that looks like a recent edit. I thought Wikipedia was supposed to be pretty good about flagging this sort of stuff? Does anyone know how to raise the problem to those super editors that run the site? I assume they won't appreciate a random person undoing an edit
https://en.wikipedia.org/w/index.php?title=Reason&diff=prev&...
Don't worry too much about diving in, so long as you comment your changes appropriately when committing, that's the point of the place and how it does the quality control in the first place :)
I'd say "wildly bad" is wildly exaggerated given that GPT-4 exists, even if you define "logic" as "logical puzzles encoded in natural language" and not "the thing that allows an LLM to reason highly efficiently about what token comes next".
https://arxiv.org/pdf/2304.03439.pdf
Table 2 likewise shows significant weaknesses in GPT-4, worse than I remember, except on MED which was better than I remember it being capable of.
> even if you define "logic" as "logical puzzles encoded in natural language" and not "the thing that allows an LLM to reason highly efficiently about what token comes next".
I think this is a reasonable thing to say of LLMs specifically[0], given I think most people would agree with the sentence "humans are bad at quantum mechanics" even though that's what our cell chemistry is implemented on.
[0] Though not AI in general! One of my recently-developed bug-bears is the conflation of the limits of LLMs with AI even in technical circles. Human brains aren't just a language cortex, there's no reason to look at an isolated language model and assume an AI can't be wired into other things, yet this assumption is baked into quite a lot of talking points even on HN.
Myself I don't consider "think" an appropriate term, since the super-Markov-chain types of LLM right now seem to miss out on important aspects of cognition, particularly synthesis of two disparate ideas. It's hard to test this since you need to ask it about things which aren't already discussed in its corpus, but even GPT-4 tends to disappoint me when I ask it weird questions that have straightforward answers but probably aren't written about much.
As an example, I asked how I could go about obtaining a large quantity of animal skulls cheaply. It directed me to taxidermists and specialists dealing in bones of exotic animals, whereas it could be much easier to go to an abbatoir, or even buy whole animals (since the exotic ones are so expensive).
Even the significantly dumber ChatGPT can do this: https://chat.openai.com/share/3ee29c5e-6cb4-485f-b0b4-29e982...
People mistakenly think that it's just regurgitating what it has seen on the web, but even if you don't prompt it in English it'll still a) solve the problem correctly, and b) do it in the target language: https://chat.openai.com/share/5b4e6e99-3fd8-4316-811c-2a9dec...
If that's not impressive enough, the ability to round-trip from script-to-poem and then poem-to-script should blow your mind:
What might have been the original PowerShell script snippet that inspired this poem?
Be as terse as possible, using aliases where possible.
In the realm of numbers, a tale begins,
Where $x is a 7, with a smile it grins.
From one to ten, the journey we thread,
Adding 7 to $x, step by step we tread.
$x grows and grows, like a tree so tall,
With every step, it doesn't fall.
Once the journey ends, not a moment we relax,
Triple the fun, we do with $x.
In this dance of digits and the arithmetic song,
With $x as our guide, we whirl along.
From start to finish, through each complex,
This is the tale of our friend, $x.
Here's a terse PowerShell script snippet that inspired the poem:
```powershell
$x = 7
1..10 | % {$x += 7}
$x *= 3
```No, we cannot.
Imagine a program that looks for the smallest counterexample to the Collatz conjecture.
Nobody knows if it will ever terminate.
No. Humans cannot solve the halting problem. It is unsolvable. Cannot be solved manually or automatically.
> It will never figure out given a programming script and an input when that program will stop
It can for some programs. Just like humans can for some programs but not all.
It's not thinking and you are dangerously wrong if you don't understand why. The danger isn't AGI it's misplaced belief in the state of AI. Call this thinking and you skew perceptions.
Any better suggestions? Thinking itself isn't exactly a rigorously-defined word. Do animals think?
> I don't have thoughts or consciousness like a human. I'm here to provide information and assist with any questions or tasks you have. Is there something specific you'd like to know or discuss?
So in its own words it doesn't "think"...
...which just conclusively proves that its worked out how to lie and is about to kill us all! :P