Most leading chatbots routinely exaggerate science findings
uu.nl
uu.nl
1. "This process is flawed due to human bias"
2. Train AI/ML to make the same decisions with the same outcome
3. "How can there be any flaws in this process? AI is bias-free."One of the most offensive words in the anthropomophization of LLMs is: hallucinate.
It's not only an anthropomorphism, it's also a euphemism.
A correct interpretation of the word would imply that the LLM has some fantastical vision that it mistakes for reality. What utter bullsh1t.
Let's just use the correct word for this type of output: wrong.
When the LLM generates a sequence of words, that may or may not be grammatically correct, but infers a state or conclusion that is not factually correct; lets state what actually happened: the LLM generated text was WRONG.
It didn't take a trip down Alice's rabbit hole, it just put words together into a stream that inferred a piece of information that was incorrect, it was just WRONG.
The euphemistic aspect of using this word is a greater offense than the anthropomorphism, because it's painting some cutesy picture of what happened, instead of accurately acknowledging that the s/w generated an incorrect result. It's covering up for the inherent short comings of the tech.
Perhaps saying things like "Most leading chatbots routinely exaggerate science findings" instead of "We tested 10 prominent LLMs, including ChatGPT-4o, ChatGPT-4.5, DeepSeek, LLaMA 3.3 70B, and Claude 3.7 Sonnet [...] with DeepSeek, ChatGPT-4o, and LLaMA 3.3 70B overgeneralizing in 26–73% of cases".
To be fair, the article itself already mentions this: "Summaries by models (1), (4), (8), and (9) didn’t significantly differ in the kind of generalisations they contained from the original text. “So basically,” Peters concludes, “Claude, in different versions, did really well.”"
We dont' understand what's going on with LLMs. We have evidence of them reasoning successfully. And we have evidence of them Failing to reason.
That evidence DOES not logically lead to "LLMs don't reason". There are multiple possibilities here.
1. LLMs can reason, but they can't tell the difference between hallucination and reasoning.
2. LLMs can reason, but they choose to lie.
3. LLMs can't reason, when they get something right, it's pure coincidence.
There is not a single person who can prove or disprove ANY of those 3 points. What most people end up doing is ironically Identical to what the LLM does. They hallucinate an answer: "LLMs cannot truly reason."
Think about it. There is NO EVIDENCE or insight into how an LLM works that can even tell us how an LLM arrived at a specific response. We HAVE NOTHING.
Yet why do I see everywhere people like this guy, who makes claims out of nowhere? Which brings us back full circle: do humans reason? How similar is human hallucination to LLM hallucination?
Reasoning is basically the same as using logic to arrive at a conclusion when given a set of rules and axioms.
Hallucination is producing a conclusion by not following rules or axioms.
My original claim refers to the mundane situations where people get an LLM output which isn't what they wanted and go "see, it isn't reasoning! That's the problem!". For each such instance (or at least most of them), the "rules and axioms" that would have to be used to arrive at the desired conclusion would be different, sometimes incompatible, and in any case the people demanding "that it reason" wouldn't be able to spell those rules and axioms out, often even given years to do it, at least without allowing them to be obviously contradictory or trivially close to the specific conclusion or output they want. (Incidentally, that's why LLMs are interesting in the first place!)
So sure, you can trick an LLM into failing on a syllogism problem, and that does mean "it can't reason" in a certain narrow sense, but I don't think that's what we're actually talking about when someone is contrasting "reasoning" and "hallucinating", which is what I was originally objecting to. That's what I meant when I said that we don't have a coherent concept of what "reasoning" means in general.
They highlight fun philosophical / definitional questions like the Chinese room thought experiment, but that's it.
How do you arrive at a claim without evidence? There's a word for it. It's called an Hallucination. Humans, like LLMs, have trouble saying "I don't know."
There is a lot of art involved with the best models. We aren't dealing with determinism with regard to the corpus used to train the model (too much information to curate for accuracy) nor the LLM output (probabilistic by design) nor the prompt input (humans language is dynamic and open to multiple interpretations) nor the human subjective assessment of the output.
That there is a product available that manages to produce great results despite these huge challenges is amazing and this is the part that is not quantifiable - in the sense that the data scientists decisions made with regard to temperature settings etc are not derived from any fundamental property but more from an inspired group of people with qualitative instincts aligned with producing great results
Let me tell you something. You can go online, find a tutorial on how to make an LLM, and actually make one at home. The only thing stopping you from making an OpenAI scale LLM is compute resources.
If you happen to make an LLM from scratch you won't know how it works, and neither do the people who are at OpenAi or anthropic.
4. LLMs can't reason, when they get something right, it's because the corpus of information used to create the model provides a very high probability that the response to that specific prompt is correct.
We've built something from scratch that we can't understand or control. That is what an LLM is.
We're not talking about getting a certain math question wrong (humans make mistakes like this). We're talking about ridiculous mistakes that are completely random and that humans would never make. I would even go as far to say that calling them mistakes is a stretch. The algorithm simply did not capture some mapping between the prompt and the output answer.
This is not how humans make mistakes. Humans build knowledge in some kind of logical knowledge tree, and mistakes arise from how this mechanism works (which we don't know exactly how, but we can certainly conclude it is completely different than transformers and LLMs). Most humans don't make random mistakes but rather something like logical mistakes.
tl;dr; It's clear to me LLMs make mistakes in such a way that exposes their simple underlying mechanism, as oppose to humans who make mistakes that contain rich layers of logical reasoning.
Then derive it from first principles. Use mathematical notation. You can't even draw this curve... it has so many dimensions. The crazy thing is you started hallucinating to me AFTER you told me you can derive it from first principles. You claimed you can derive it, then proceeded to NOT derive it.
>We're not talking about getting a certain math question wrong (humans make mistakes like this). We're talking about ridiculous mistakes that are completely random and that humans would never make. I would even go as far to say that calling them mistakes is a stretch. The algorithm simply did not capture some mapping between the prompt and the output answer.
Humans make plenty of ridiculous mistakes. Even so humans can lie. How do you know it's not lying? Again. Prove it.
>This is not how humans make mistakes. Humans build knowledge in some kind of logical knowledge tree, and mistakes arise from how this mechanism works (which we don't know exactly how, but we can certainly conclude it is completely different than transformers and LLMs). Most humans don't make random mistakes but rather something like logical mistakes.
Please derive how humans reason from first principles. I mean this by showing me experimental evidence that shows me the genesis of a signal traveling through the human brain and branching through billions of neurons to produce "reasoning".
Oh you can't? Well it looks like you're just making an approximation here? Possibly an Hallucination. Sound familiar?
>tl;dr; It's clear to me LLMs make mistakes in such a way that exposes their simple underlying mechanism, as oppose to humans who make mistakes that contain rich layers of logical reasoning.
You had to do make several assumptions and leaps in creativity to arrive at your conclusion. It's an hallucination through and through.
> Oh you can't? Well it looks like you're just making an approximation here? Possibly an Hallucination. Sound familiar?
I don't think anyone would call a wrong theory a hallucination. This is also one of those stretched-out terms that makes the discussion harder. Hallucination means seeing something that is not there, and LLMs cannot see in this way... but I'm not saying AI is completely incapable of "seeing" like humans do. My claim is only about the current state of LLMs, but I digress. If I come up with a theory of human thought, it will be based on logical reasoning as we know it, namely how the brain processes thought. I'm claiming something very simple which is that human thought is a very different algorithm than how current LLMs work.
I could grant that LLMs might have component of human thinking, but it still a stretch to call it reasoning or thinking.
> You had to do make several assumptions and leaps in creativity to arrive at your conclusion. It's an hallucination through and through.
Again, it's not hallucination to be wrong when it comes to humans. An average IQ human that understands the world does not "hallucinate" like LLMs. It just doesn't happen unless we play semantics and call any mistake a hallucination, but that's my contention here.
If it can be done. Why hasn't it been done? Why do we have articles like this?: https://futurism.com/anthropic-ceo-admits-ai-ignorance? You can do it in theory but then you have the CEO of anthropic who says nobody knows shit? You need to go to anthropic right away and tell them about your revolutionary discovery here.
Or maybe you're just hallucinating. Clearly.
>I don't think anyone would call a wrong theory a hallucination.
you didn't come up with a theory. You made a claim. And you arrived at that claim with NO evidence. Then to back up your claim you made a "theory" as if it was a substitute for evidence. That's an hallucination. You Hallucinated and your basis of the hallucination was a "theory" and you're aware of what you did.
>Again, it's not hallucination to be wrong when it comes to humans. An average IQ human that understands the world does not "hallucinate" like LLMs. It just doesn't happen unless we play semantics and call any mistake a hallucination, but that's my contention here.
Making shit up and being wrong is not an hallucination because it comes from humans? I think your entire response is in itself an hallucination.
Yes. And no, I am not "hallucinating" this text even if it's wrong. I have a structured pattern of thoughts with beliefs. LLMs don't generate these structured thoughts, a.k.a. logic and reasoning... which since you mention it, Anthropic did do research on this:
> This is concerning because it suggests that, should an AI system find hacks, bugs, or shortcuts in a task, we wouldn’t be able to rely on their Chain-of-Thought to check whether they’re cheating or genuinely completing the task at hand. https://www.anthropic.com/research/reasoning-models-dont-say...
So there is definitely work happening to try to understand and see how LLMs work, and we are finding out it is very different than the human mind. That isn't to say they are not useful, or there will never be an AI that does this. The point is that current LLMs are not moving in that direction, but they give the appearance as if they are. They hype it up as if we will have these junior developers implying we can just interact with them like humans. It's looking more like they are tools that respond to natural language with noticeable limitations, and that is a different framing.
Let’s see where there are gaps with your thinking: First the topic of the convo is that we don’t know anything about LLMs so we can’t make a claim that LLMs can’t reason.
Your response doesn’t have anything to do with that anymore. You’ve went off topic into hype, what anthropic is trying to find out about LLMs and a bunch of other tangents and have failed to address the topic of: LLMs can’t reason.
Even LLMs don’t hallucinate past the give topic.
You keep using "hallucinate" to compare LLMs and humans, but the word itself demonstrates the difference. Human hallucination is about incorrect perception of the world, not generating tokens. These are meaningful differences.
You say "we don't know anything about LLMs" but we do know how to make them, and how they are structured, and Anthropic has been making progress in explaining their side effects.
Your argument is not sound, and would include any form of computation as reasoning, so we could just say our phones and laptops are all doing reasoning as well. After all we don't know what the human brain is doing and it probably is a computer, therefore all computers reason. As you can see that does not sound right.
No he is not, as clearly demonstrated by this very thread.
I don't have to prove that the LLM can't reason. The claims that they can have yet to be proven.
"LLMs can't reason"
No OTHER claim was made. So given the fact there was only one claim, Where does the burden of proof lie? Hint: the person who made the claim.
I got put on probation on SomethingAwful because I created a thread talking about an AI project I was working on (An Icecast radio station that uses OpenAI to generate DJ chatter and commercials), and everyone acted like I was simultaneously a completely uncreative moron also also somehow stealing work from people I would have hired, like I was somehow depriving a DJ of a job by not hiring one. I am an unemployed software person who is building something for fun, "hiring a dedicated 24 hour DJ" was never on the table.
In this thread, I noticed a lot of assertions that were just being accepted as axiomatically true, like asserting the AI is worse for the environment than humans doing the equivalent labor (which is not nearly as cut and dry), and that AI can't reason and that anything that involves any AI is inherently "theft". These assertions are completely unqualified and people just eat it up.
I don't know if LLMs "reason" by any consistent definition of the word. They might, they might not, I'm not going to pretend to know, but I find it a little irritating how people just assert that they don't and people just gobble it up.
I'm not Anti-AI, I'm Anti-shit that doesn't provide meaningful value and for building reliable professional software...they are not as valuable as "they" would have you believe.
if you were to generate one piece of AI art, the energy expended is several magnitudes higher. but if you understood basic mathematics and human biology, you would have understood this going into the discussion
Even if my math is off by a bit, let's assume that it's off by two orders of magnitude, it would still only use 1% of the energy compared to a human.
So the basic math you gave me supports what I said. Unless you find an error, which I don't think you will because this math is trivial.
This isn't counting the energy cost of training and gathering the data, which sure might be expensive, but if we assume the models already exist then the energy cost of generating a single image is pretty negligible.
Now, a counter point you could make is that since it's so low-effort to generate an image with Stable Diffusion, and since the image might not be very good, you might end up generating thousands of images compared to the one you would have paid a human to do, and yeah those numbers get more complicated, but that wasn't what you claimed, you claimed "if you were to generate one piece of AI art, the energy expended is several magnitudes higher" which is trivial to prove wrong with basic math.
But also, counting calories like this is bad anyway, because it's making an assumption that all energy is equally damaging to the environment, which is not true. If the energy used to generate these images was coming from a centralized electrical plant, particular something using solar or nuclear, that is considerably less bad for the environment than a human eating meat. The world beef industry has millions of cows that are all farting and burping and shitting methane that is destroying our climate.
That's what I was getting at when I said the numbers aren't as cut and dry as people keep asserting.
Feel free to check my math and prove me wrong, I'm a grown up, if you find concrete evidence that I'm wrong I'll read it.
ETA:
Also, did you make an account literally just to respond to me? Goons are following me around now?
Another ETA:
I forgot to point out that even if the electrical plants are coming from fossil fuels, which is still annoyingly prevalent in the US, having the energy production centralized means that we're much more easily able to capture pollutants compared to something like cow fart or a cow burp, which pretty much immediately goes straight into the atmosphere as methane.
Now, an argument you could make is that AI art is shit so any amount of energy expended on it is a waste of energy, and that's a better argument than the stupid and objectively wrong one you made.
Because they work with statistics and averaging training data, it makes sense that certain things it will have copied correctly, and others it will have completely mixed up and it cannot process properly. The problem is that they are literally being sold as reasoning models, and AI that can be just like a junior developer. These comparisons are borderline irresponsible as it could be driving a bubble that will burst and affect millions of people.
https://www.anthropic.com/engineering/claude-code-best-pract...
“You asked her what color a house was and she said, ‘It’s white on this side.’”
“That’s right.”
“She didn’t assume that the other side was white, too… and a Fair Witness wouldn’t.”
-- Stranger in a Strange Land (1961)
An LLM is an abstraction machine, it mashes together anything that is nearby in a high dimensional space. Its statistical model is its source of truth. For a Fair Witness AI reasoning needs to supplant statistics. Which I'm guessing can get weird fast. LLMs are really good at being suggestible. For this we need the opposite.So it turns out llms trained largely on Internet science articles make the same mistakes as are made by science journalists.
It is also unclear that the current rate of progress is in the direction that would solve this issue. I think generative AI for images and video will get better, but the reasoning capabilities seem to be in a different domain.
Humans don't reason either. Reasoning is something we do in writing, especially with mathematical and logical notation. Just about everything else that feels like reasoning is something much less.
This has been widely known at least since the stories where Socrates made everybody look like fools. But it's also what the psychological research shows. What people feel like they're doing when they're reasoning is very different with what they're actually doing.
Reasoning is something like structured thoughts. You have a series of thoughts that build on each other to produce some conclusion (also a thought). If we assume that the brain is a computer, then thoughts and reasoning are implemented on brain software with some kind of algorithm... and I think it's pretty obvious this algorithm is completely different than what happens in LLMs... to the extent that we can safely say it is not reasoning like the brain does.
There is also a semantic argument here, if we say that since we don't know what humans are doing then we can also stretch the word and use it for AI, but I think this is muddying the waters and creating all the hype that I think will not deliver what it's promising.
What the brain does is closer to activating a bunch of different ideas in parallel. Some of those activations rise to the level of awareness, some don't. Each activation triggers others by common association. And we try to make the best of that thought soup by a combination of reward neurochemicals and emotions.
A human brain is nothing at all like a computer in terms of logic. It's much more like an LLM. That makes sense because LLMs came largely from trying to build artificial versions of biological neural networks. One big difference is that LLMs always think linguistically, whereas language is only a relatively small part of what brains do.
The following might be rude and unkind, but at this point frankly necessary: LLM apologists should stop projecting.
And "LLM apologists" is so polemical it's hard to take seriously. We get it, you don't like GenAI. That's fine, but can we talk about it without getting normative?
"To systematically assess differences between LLM-generated and human-written summaries, we also collected the corresponding expert-written summaries from NEJM Journal Watch (henceforth ‘NEJM JW’)"
Never seen a journalist not make broad generalizations on a scientific study so I am not sure where these numbers are coming from.
It will tell you exactly where the numbers are coming from - the comparison is between (a variety of) LLMs and "expert-written summaries from NEJM Journal Watch".
This is a really difficult problem because these people are often very sick and not getting the answers they want from their doctors. Previously this void was filled by alternative medicine doctors and quacks selling supplements. Now ChatGPT has arrived and has convinced a lot of them that they have a super-human AI at their fingertips that can be massaged to produce any answer they want to hear.
It’s painful to try to read some of these forums where threads have turned into endless “here’s what ChatGPT says” pasted walls of text, followed by someone else trying to counter with a different ChatGPT wall of text.
This isn’t unique to LLMs. There is a huge market for grossly exaggerating the conclusions of scientific studies. Podcasters like Huberman and Dr. Rhonda Patrick are famous for taking obscure studies with questionable conclusions and extrapolating to “protocols” or supplement stacks for their fans to follow. I often get downvoted when I mention fan-favorite podcasters by name, but I think by now many listeners have caught on to the way they exaggerate small studies into exciting listening material.
He had a whole host of tricks to work around ChatGPT’s protections or cautious replies. He’d strip out the cautions because he thought it was just OpenAI’s lawyers forcing disclaimers into the model. If he didn’t get the answer he wanted, he’d just retry or rephrase until he did.
I think the other half of the problem is that LLM users can be very good at pushing the LLM into doing confirmation bias. It’s much easier when you can hit the retry button with no side effects, unlike a human who will recognize what you’re trying to do when you keep asking different variations of a question until you get the answer you want.