> I hope you can see the stupidity here if you expect to see any deterministic results at all.
Are you expecting humans to be deterministic in the code they produce?
> I hope you can see the stupidity here if you expect to see any deterministic results at all.
Are you expecting humans to be deterministic in the code they produce?
And?
The p(that kind of error) is pretty small now. At what point does a probability coming out of an LLM look like "knowing", such that spitting out the wrong answer despite that probability looks like a health problem, a typo, or even just boredom? (Thinking of the Lizardman constant here: https://en.wiktionary.org/wiki/Lizardman%27s_Constant)
It's a continuum for both them and us, even if the mechanism is wildly different.
> Making mistakes is not the same as non-deterministic.
i.e. when the dismissal is "non-deterministic" when it should be "Making mistakes", is itself a mistake.
A human who knows 1+1=2 can still say “3” because they misread the question, misspoke, were distracted, or made some other cognitive error. Likewise, an LLM can output “3” because the generation process selected an incorrect continuation. Those are both errors in producing an answer, not evidence that 1+1 somehow has multiple answers.
So yes, human mistakes and LLM sampling are mechanistically different. If your argument is that LLMs and humans can both make mistakes, then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans.
It's not, I'm just pointing out that LLMs won't make that mistake.
You could ask an LLM what 1+1 is, and the number of times it says "3" is so small that it makes no sense to worry about it. It will phrase the response differently each time; that's the nondeterminism. But it won't say "3".
> then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans.
Yes, if we ignore everything else, that seems like a reasonable question. But let's not ignore everything else, like the fact that LLMs are much more productive than humans and likely already make fewer mistakes than the average programmer.
I think the disturbing fact is that you can take a frontier model with all the intelligence of humanity, and make it say 1 + 1 = 3, by specifically training for it...
A human with that much knowledge will refuse that attempt. There in lies the difference..
And you seem confused about how an LLM works and what it is - “the intelligence of all humanity” - not how it works. You’re debating using 2 imaginary things you created.
Hit them with a stick until they answer as you told them to.
I think the post up-thread, https://news.ycombinator.com/item?id=49881653, was trying to make this point by linking to an episode of Star Trek TNG, with Picard being tortured until he said the "correct" (incorrect) number of lights.
(I recommend against using fiction as evidence; in this case the general point happens to be valid, and is why torture is forbidden: we humans really do break, but breaking doesn't mean we tell the truth, it means we tell people what we think they want to hear).
Obviously, for this purpose, human should not have any feelings (because LLMs don't have), so can't feel pain. Or else the comparison can't work.
Other than this forcing you to ignore the overwhelming majority of humans who have functioning pain nerves:
LLMs have something functionally equivalent to pain, in this regard at least.
During training, model weights are updated depending on if the feedback was positive or negative.
It has a functional effect similar to that which pleasure and pain have with us. Not identical, so far as I know there's not been any reports of any machine learning model that is into BDSM, but for the most part functionally similar.
There is also research which has found circuits in multiple LLM models, which serve similar roles at inference time and are distinct from other emotional representations: https://arxiv.org/abs/2609.16247
Well, that's not true.
https://www.youtube.com/playlist?list=PLO3a3Ax6Yh6bbtKuxfYBP...
But this is plainly false. This kind of unforced error occurs all the time.
For example, once when I was in high school I traced an error in my math homework to an intermediate calculation of "2 + 2" as being "3". There was no reason.
What we can say about humans is that, if they know that 1 + 1 = 2, (a) they are unlikely to change their mind about this in any kind of lasting or permanent way, and (b) the rate at which they will mistakenly produce other values for 1 + 1 is very low. But it will happen occasionally, and when it does happen, "they just suddenly decided on the wrong value" is an extremely accurate description of what that looks like.
The behaviour/output of an LLM is not like that. Ask an LLM to create a dashboard to show games by genre and it will generate different results with each run, and each model/model version produces wildly different results.
As humans we don't have our memory reset multiple times per day
I've seen humans vote for Brexit, re-elect Trump, ask questions clearly already answered in an FAQ, try to pull on a door labelled "push", and insist on giving me homeopathic silicon dioxide pills* that cost £5** for a 10-12 gram packet.
Continual learning is a difference, but not by itself a reason to care about "deterministic results".
Nor, indeed, correct results.
> As humans we don't have our memory reset multiple times per day
Humans need sleep well before they can read a million tokens' worth of written text. We're more like 300k tokens if you're actually reading and not skimming for 16 hours straight.
Again, different (in soooo many ways), but this isn't a relevant difference when the topic is "deterministic results".
* yes, sand: https://dailymed.nlm.nih.gov/dailymed/fda/fdaDrugXsl.cfm?set...
** and that was what it cost in the 90s
There is no RNG involved when I decide to push vs pull the unlabeled door to my building every morning, it becomes deterministic because its baked into memory
You can put stuff in context to deal with this but you can't do that for everything, its not practical and you would blow the context window
You, me, everyone, we all fail tasks with some non-zero probability, including tasks we've done many times before. This is literally why typos happen. We don't always catch the typos we make, this is literally why spellcheck exists. We don't always catch logical errors etc., this is literally why professional writers work with professional editors. We don't always spot errors in code we write; this was why compiler errors have to exist at one level, why unit tests and integration tests exist at another, and why despite automated tests we have QA departments.
Statistically speaking, you as a human will have made more comprehension errors reading that paragraph than even LLMs a few years back.
Not zero, just fewer. They're not magic and neither are we.
> There is no RNG involved when I decide to push vs pull the unlabeled door to my building every morning, it becomes deterministic because its baked into memory
There's a lot of random inputs into human behaviour. How much they become random outputs is a similar question as for AI. But like I said, typos are a trivial existence proof of humans making randomised errors on a thing they've done many times before. Timing error at random and a "the" becomes a "teh".
Here's one for gross failure modes being random: Small (and apparently random) groups of neurons in your (and all animal) brains briefly enter sleep-like states even while you remain apparently awake. This starts well before you even feel tired. And like I said, looks to be random.
Your memories are not "baked" at any point; they're made of and by cells doing chemistry. Your entire body is a trillion small bags of chemicals that communicate with each other by leaking out signalling molecules and electrochemical gradients into each other's water supply, and this goes wrong sometimes. Human memory is known to be fallible, and also manipulable after the fact.
In the sense that if I hire a competent person, I can tell them something and they will deterministically do it
For an LLM I have no mechanism to even update the weights
All you can do is play around with context, which is like taping a post-it note to someone's desk
Imagine you had an employee where you had to tape 1000 post-it notes to their desk