Ultimately, an LLM models language and the process behind it's creation to some degree of accuracy or another. If that model includes a way to approximate the act of reasoning, then it is reasoning to some extent. The extent I am happy to agree is open for discussion, but that reasoning is taking place at all is a little harder to attack.
This would form a filter not unlike, yet distinct from, our understanding of personal experience.
you could make the exact same argument against humans, we just learn to make sounds that elicit favourable responses. Besides, they have plenty "skin in the game", about the same as you or I.
If you want to understand where a systems limitations are you need to understand not just what it does but how it does it, I feel like we need to start teaching classes on Behaviorism again.
Are you sure you aren’t just defining reasoning as something only a human can do?
You know how we insert ourselves into the process of coming up with a delicious recipe? That first person perspective might be also necessary for reasoning. No computer knows the taste of mint, it must be given parameters about it. So if a computer comes up with a recipe with mint, we know it wasn’t via tasting anything ever.
A calculator doesn’t reason. A facsimile of something we have no idea about its role in consciousness has the same outlook as the calculator.
Jet planes are so far removed mechanically from a bird that the idea they fly is not even remotely worth considering.
But since we knew properties of wings were major comments to flight dating back to beyond the myths of Pegasus or Icarus, we rightly connected the similarities in the flight case.
Yet while we have studied neurons and know the brain is apart of consciousness, we don’t know their role in consciousness like the wing’s for flight.
If you got a bunch if daisy chained brains and that started doing what LLMs do, I’d change my tune—because the physical substrates are now similar enough. Focusing on neurons, and their facsimilized abstractions, may be like thinking flight depending upon the local cellular structure of a wing, rather than the overall capability to generate lift, or any other false correlation.
Just because an LLM and a brain get to the same answer, doesn’t mean they got there the same way.
Bailey? Reason.
How reasonable are the outputs of ANNs considering the inputs? This is a valid question and it has a useful response.
From ImageNet to LLMs we are finding these tools to give some scale of a reasonable response.
Recommended reading: Philosophical Investigations by Wittgenstein.
If not, then why shouldn’t differently constructed but algorithmically similar systems be able to produce similar phenomena?
See C-sections versus natural birth. Formula versus mother's milk. Etc.
What non-sentience based property do you think something should have to be considered reasoning. Do you consider sentience and reasoning to be one and the same? If not then you should be able to indicate what distinguishes one from the other.
I doubt anyone here is arguing that chatGPT is sentient, yet plenty accept that it can reason to some extent.
No, but I think they share some similarities. You can be sentient without doing any reasoning, just through experience, there's probably a lot of simple life forms in that category. Where they overlap I think, is in that they require a degree of reflection. Reasoning I'd say is the capacity to distinguish between truth and falsehoods, to have mental content of the object you're reasoning about and as a consequence have a notion of understanding and an interior or subjective view.
The distinction I'd make is that calculation or memorization is not reasoning at all. My TI-83 or Stockfish can calculate math or chess but they have no notion of math or chess, they're basically Chinese rooms, they just perform mechanical operations. They can appear as if they reason, even a chess engine purely looking up results in a table base and with very simplistic brute force can play very strong chess but it doesn't know anything about chess. And with the LLMs you need to be careful because the "large" part does a lot of work. They often can sound like they reason but when they have to explain their reasoning they'll start to make up obvious falsehoods or contradictions. A good benchmark if something can reason is probably if it can.. reason about its reasoning coherently.
I do think the very new chain-of-thought models are more of a step into that direction, the further you get away from relying on data the more likely you're building something that reasons but we're probably very early into systems like that.
Real reasoning is being able to manipulate symbolic expressions in a consistent manner while preserving some invariants.
Personal experience as logic is how you end up with the Holocaust.
Correct answer: https://chatgpt.com/share/67a9500b-2360-8007-b70e-0bc2b84bc1...
Incorrect answer (I think): https://chatgpt.com/share/67a950df-d4e0-8007-8105-95a9e5be19...
They aren't terrible, and they have all of arXiv to train on. Terrence Tao is doing some cool stuff with it - the idea will be an LLM to generate Lean proofs.
https://mathstodon.xyz/@tao/113132502735585408
https://terrytao.wordpress.com/2024/12/05/ai-for-math-fund/
(Professor Tao is probably the best or at least most productive in the most fields current mathematician).
Here's some cool math I learned from a regular book, not an LLM:
If a synthetic "mimicry" can displace human thinking, we've got serious problems, regardless of whether or not you believe that it's "real".
ie, that logically consistent thinking starts from symbolization, axioms, proof procedures, world models. otherwise, you end up with persuasive words.
The beautiful thing about reasoning models is that there is no need to overcomplicate it with all the things you've mentioned, you can literally read the model's reasoning and decide for yourself if it's bullshit or not.
Humans rely in cues to tell when each other is fabricating or lying. Machines don't have those cues, and fabricate their reasoning too. So we have a complicatedly difficult time trusting them.
After a conversation with humans I think you'd agree 1% of them being able to reason deeply is a vast overestimation.
A good example to see how little people can reason is the following classic:
> Given the following premises derive a conclusion about your poems:
> 1) No interesting poems are unpopular among people of real taste.
> 2) No modern poetry is free from affectation.
> 3) All your poems are on the subject of soap bubbles.
> 4) No affected poetry is popular among people of taste.
> 5) Only a modern poem would be on the subject of soap bubbles.
The average person on the street won't even know where to start, the average philosophy student will fuck up the translation to first order logic, and a logic professor would need a proof assistant to get it right consistently.
Meanwhile o3-mini in 10 seconds:
We can derive a conclusion about your poems by following the logical implications of the given premises. Let’s rephrase each premise into a more formal form:
Premise 1: No interesting poems are unpopular among people of real taste. This can be reworded as: If a poem is interesting, then it is popular among people of real taste.
Premise 2: No modern poetry is free from affectation. This tells us: If a poem is modern, then it is affected (i.e., it shows affectation).
Premise 3: All your poems are on the subject of soap bubbles. In other words: Every one of your poems is about soap bubbles.
Premise 4: No affected poetry is popular among people of taste. This implies: If a poem is affected, then it is not popular among people of taste.
Premise 5: Only a modern poem would be on the subject of soap bubbles. This means: If a poem is about soap bubbles, then it is modern.
Now, let’s connect the dots step by step:
From Premise 3 and Premise 5:
All your poems are on the subject of soap bubbles.
Only modern poems can be about soap bubbles.
Conclusion: All your poems are modern.
From the conclusion above and Premise 2:
Since your poems are modern, and all modern poems are affected,
Conclusion: All your poems are affected.
From the conclusion above and Premise 4:
Since your poems are affected, and no affected poem is popular among people of taste,
Conclusion: Your poems are not popular among people of taste.
From Premise 1:
If a poem is interesting, it must be popular among people of taste.
Since your poems are not popular among people of taste (from step 3), it follows that:
Conclusion: Your poems cannot be interesting.
Final Conclusion: Your poems are not interesting.
Thus, by logically combining the premises, we conclude that your poems are not interesting.
Except, human mimicry of "reasoning" is usually applied in service of justifying an emotional feeling, arguably even less reliable than the non-feeling machine.
LLMs? I'm waiting for one that knows how not to say something that is clearly wrong with extreme confidence, reasoning or not.
1) understanding anything == building a causal model of it
2) intelligence == ability to build causal models
3) reasoning == proving or disproving statements
4) math == causal models of abstract worlds
5) science == causal models of real world with associated real world actions to test hypothesis
over time this question has been debated by philosophers, scientists, and anyone who wanted to have better cognition in general.
You might want to brush up on your Greek history.
Now if it quacks like a duck in 95% of cases, who cares if it's not really a duck? But Google still claims that water isn't frozen at 32 degrees Fahrenheit, so I don't think we're there yet.
Somehow it always seems to end up at eugenics and white supremacy for those people.
llm, meanwhile, is putting out plausible tokens which is consistent with its training set.