Statistical or probabilistic reasoning is still reasoning, and indeed essentially mandatory to make sense of the real world, and is what humans do. Indeed the failure of painstakingly hand-trained symbolic reasoning engines to model reality (as opposed to just simple toy worlds) was the entire reason for the rise of neural networks.
The entire point of a neural network is to efficiently compress large amounts of statistical data into a model with as much predictive power as possible. Somewhere within GPT-4's many layers there has to exist something like Bayes networks and other ultra-efficient representations of the vast network of correlations present in the training corpus – otherwise there's no way it could do what it does.
I didn't mean this by "reasoning". To me, Coq does not reason (in the meaning I intended). It checks proofs are well-formed and assists proof writing.
But yes, there's some implementation of logic in Coq. It applies (a specific form of) logic. Again, LLMs output probable text. They don't do logic. They do statistics on language.
Coq works on the actual objects themselves directly, not on the text that presents them.
LLMs usually fail to answer to riddles correctly for this reason. We saw plenty of occurrences of this in recent times. Each time LLMs are benchmarked, people think of trying riddles and it works to some extent because something that looks like these riddles are in the training set but it eventually fails. Humans also make mistakes when solving riddles, but not for the same reasons (the process is not the same, they are not "processing text", the are really processing logic, even if the process is fallible).
You do not know what humans do, no human does, as we don't have such an access to our internal representations. You neither know what LLMs do, no human does, as even though we have full access to their internals, we have little to no idea how to interpret said internals.
The most interesting thing about LLMs is the degree that they have learned to model the real world while only being exposed to words. Because those words are highly correlated not just with each other, but with an external entity unknown to the LLM, it's very likely that by learning to model that external entity, the LLM's word-predicting power will increase. LLMs are likely to have converged on a world-model simply because it makes the most sense given the available evidence.
This is no different from how human scientists create models of things they cannot directly see based on what they can see. And of course, "seeing" is not directly perceiving the external reality either! Sensory data merely correlates with what's really going on, and the brain has learned to model the world – an unknown external entity – based on the correlations that exist in the sensory language.
I'll admit I did this. I tried to frame / convey my perspective on the word "think".
> Not a very convincing argument.
I tried to detail a bit more though.
> You do not know what humans do
No, but I'm pretty sure I don't try to predict the next word when solving a logical problem through statistical text processing. I would probably do something more like this if I were to bullshit someone or give an insipid political speech, but I think I'm quite bad at it.
(nothing to answer on the rest of your comment, it's interesting and I quite agree)
(edit: but see https://news.ycombinator.com/item?id=37912319)
Neither do transformer / deep learning LLMs. They build an internal representation of the text that captures various relationships and structures within the broader context. They are not Markov chains.
Coq parses the text to an abstract representation and so do LLMs.
Compilers translate their abstract interpretation to another programming language, usually native assembly code or bytecode. Coq applies / verifies logical (inductive) inferences on logical objects (rules, proofs, axioms, predicates) (not a Coq expert, details should be checked). LLMs predict next words.
That's an oversimplification. Deep learning LLMs such as ChatGPT are designed to solve natural language processing (NLP) tasks, not just next-word prediction. In order to "predict next words" they create a model and "reason" about it. They are not simple Markov processes. There's a lot happening in that ~100 layer neural network.
It just turns out that good next-word prediction requires understanding the world.
(And you can equally view humans as "next word predictors" anyway.)
> Reason is the capacity of applying logic consciously by drawing conclusions from new or existing information, with the aim of seeking the truth, with more than 10 or more than 3 evidence, 99.9% reliability or 87.5% reliability.
Apart from the words "consciously" (which has 100 or so different definitions, some of which do and some of which do not apply to AI), and "aim" which can be contentious depending upon if you feel that requires consciousness or can include training rewards… I think the best of them is at the lower of those two mysterious percentages I've never seen discussed before.
And that's despite LLMs being wildly bad at, specifically, logic. Though even this specific weakness in general seems like a very easy thing to get around by hybridising it with the computer it has to run on e.g. by telling it how to use a compiler, as logic is the foundation of the arithmetic that the transformers use to do natural language comprehension in the first place.
Edit: yeah that looks like a recent edit. I thought Wikipedia was supposed to be pretty good about flagging this sort of stuff? Does anyone know how to raise the problem to those super editors that run the site? I assume they won't appreciate a random person undoing an edit
https://en.wikipedia.org/w/index.php?title=Reason&diff=prev&...
Don't worry too much about diving in, so long as you comment your changes appropriately when committing, that's the point of the place and how it does the quality control in the first place :)
I'd say "wildly bad" is wildly exaggerated given that GPT-4 exists, even if you define "logic" as "logical puzzles encoded in natural language" and not "the thing that allows an LLM to reason highly efficiently about what token comes next".
https://arxiv.org/pdf/2304.03439.pdf
Table 2 likewise shows significant weaknesses in GPT-4, worse than I remember, except on MED which was better than I remember it being capable of.
> even if you define "logic" as "logical puzzles encoded in natural language" and not "the thing that allows an LLM to reason highly efficiently about what token comes next".
I think this is a reasonable thing to say of LLMs specifically[0], given I think most people would agree with the sentence "humans are bad at quantum mechanics" even though that's what our cell chemistry is implemented on.
[0] Though not AI in general! One of my recently-developed bug-bears is the conflation of the limits of LLMs with AI even in technical circles. Human brains aren't just a language cortex, there's no reason to look at an isolated language model and assume an AI can't be wired into other things, yet this assumption is baked into quite a lot of talking points even on HN.