A human who knows 1+1=2 can still say “3” because they misread the question, misspoke, were distracted, or made some other cognitive error. Likewise, an LLM can output “3” because the generation process selected an incorrect continuation. Those are both errors in producing an answer, not evidence that 1+1 somehow has multiple answers.
So yes, human mistakes and LLM sampling are mechanistically different. If your argument is that LLMs and humans can both make mistakes, then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans.
It's not, I'm just pointing out that LLMs won't make that mistake.
You could ask an LLM what 1+1 is, and the number of times it says "3" is so small that it makes no sense to worry about it. It will phrase the response differently each time; that's the nondeterminism. But it won't say "3".
> then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans.
Yes, if we ignore everything else, that seems like a reasonable question. But let's not ignore everything else, like the fact that LLMs are much more productive than humans and likely already make fewer mistakes than the average programmer.
I think the disturbing fact is that you can take a frontier model with all the intelligence of humanity, and make it say 1 + 1 = 3, by specifically training for it...
A human with that much knowledge will refuse that attempt. There in lies the difference..
And you seem confused about how an LLM works and what it is - “the intelligence of all humanity” - not how it works. You’re debating using 2 imaginary things you created.
Hit them with a stick until they answer as you told them to.
I think the post up-thread, https://news.ycombinator.com/item?id=49881653, was trying to make this point by linking to an episode of Star Trek TNG, with Picard being tortured until he said the "correct" (incorrect) number of lights.
(I recommend against using fiction as evidence; in this case the general point happens to be valid, and is why torture is forbidden: we humans really do break, but breaking doesn't mean we tell the truth, it means we tell people what we think they want to hear).
Obviously, for this purpose, human should not have any feelings (because LLMs don't have), so can't feel pain. Or else the comparison can't work.
Other than this forcing you to ignore the overwhelming majority of humans who have functioning pain nerves:
LLMs have something functionally equivalent to pain, in this regard at least.
During training, model weights are updated depending on if the feedback was positive or negative.
It has a functional effect similar to that which pleasure and pain have with us. Not identical, so far as I know there's not been any reports of any machine learning model that is into BDSM, but for the most part functionally similar.
There is also research which has found circuits in multiple LLM models, which serve similar roles at inference time and are distinct from other emotional representations: https://arxiv.org/abs/2609.16247
Well, that's not true.
https://www.youtube.com/playlist?list=PLO3a3Ax6Yh6bbtKuxfYBP...
But this is plainly false. This kind of unforced error occurs all the time.
For example, once when I was in high school I traced an error in my math homework to an intermediate calculation of "2 + 2" as being "3". There was no reason.
What we can say about humans is that, if they know that 1 + 1 = 2, (a) they are unlikely to change their mind about this in any kind of lasting or permanent way, and (b) the rate at which they will mistakenly produce other values for 1 + 1 is very low. But it will happen occasionally, and when it does happen, "they just suddenly decided on the wrong value" is an extremely accurate description of what that looks like.
And?
The p(that kind of error) is pretty small now. At what point does a probability coming out of an LLM look like "knowing", such that spitting out the wrong answer despite that probability looks like a health problem, a typo, or even just boredom? (Thinking of the Lizardman constant here: https://en.wiktionary.org/wiki/Lizardman%27s_Constant)
It's a continuum for both them and us, even if the mechanism is wildly different.
> Making mistakes is not the same as non-deterministic.
i.e. when the dismissal is "non-deterministic" when it should be "Making mistakes", is itself a mistake.