LLMs don't understand. It's mind-boggling to me that large parts of the tech industry think that.
Don't ascribe to them what they don't have. They are fantastic at faking understanding. Don't get me wrong, for many tasks, that's good enough. But there is a fundamental limit to what all this can do. Don't get fooled into believing there isn't.
But I'm afraid that most folks using the term mean it more literally than you describe.
I think you might be tied to a definition of "understanding" that doesn't really apply.
If you prompt a LLM with ambiguous instructions, it requests you to clarify (i.e., extend prompt to provide more context) and once you do the LLM outputs something that exactly meets the goals of the initial prompt, does it count as understanding?
If it walks like a duck and quacks like a duck, it's a duck,or something so close to a duck that we'd be better off calling it that.
It does not understand that it needs clarification. This behavior is replicated pattern
Such as, trying to get an AI to create a design mockup of mastodon but it keeps falling over at laying out the page correctly: https://social.coop/@scottjenson/114593227688326501
* replying with “I don’t know” a lot more often
* consistent responses based on the accessible corpus
* far fewer errors (hallucinations)
* being able to beat Pokémon reliably and in a decent time frame without any assistance or prior knowledge about the game or gaming in general (Gemini 2.5 Pro had too much help)
Your question can be rephrased to “what would an actual difference look like.”
However, what you are asking underneath that, is a mix of “what is the difference” and “what is the PRACTICAL difference in terms of output”
Or in other words, if the output looks like what someone with understanding would say, how is it meaningfully different.
—-
Humans have a complex model of the world underlying their thinking. When I am explaining this to you, you are (hopefully) not just learning how to imitate my words. You are figuring out how to actually build a model of an LLM, that creates intuitions / predictions of its behavior.
In analogy terms, learning from this conversation, (understanding) is to create a bunch of LEGO blocks in your head, which you can then reuse and rebuild according to the rules of LEGO.
One of the intuitions is that humans can hallucinate, because they can have a version of reality in their head which they know is accurate and predicts physical reality, but they can be sick/ill and end up translating their sensory input as indicating a reality that doesn’t exist. OR they can lie.
Hallucinations are a good transition point to move back to LLMs, because LLMs cannot actually hallucinate, or lie. They are always “perceiving” their mathematical reality, and always faithfully producing outputs.
If we are to anthropomorphize it back to our starting point about “LLMs understand”, this means that even when LLMs “hallucinate” or “lie”, they are actually being faithful and honest, because they are not representing an alternate reality. They are actually precisely returning the values based on the previous values input into the system.
“LLMs understand” is misleading, and trojans in a concept of truth (therefore untruth) and other intuitions that are invalid.
—-
However, understanding this does not necessarily change how you use the LLMs 90% of the time, it just changes how you model them in your head, resulting in a higher match between observer reality and your predictive reality.
For an LLM this makes not difference, because its forecasting the next words the same way.
(Ironically, I would not be too surprised if it was produced by an LLM.)
In the first prompt the replicated pattern is to ask for clarification, in the second prompt the replicated pattern is to perform the work. The machine might understand nothing but does it matter when it responds appropriately to the different cases?
I don't really care whether it understands anything at all, I care that the machine behaves as though it did have understanding.
No. You have an initial prompt that is vague, and then you have another prompt that is more specific.
- "draw me an automobile"
- "here's a picture of an ambulance."
- "could you make it a convertible instead? Perhaps green."
- "ok, here's a picture of a jaguar e-type".
Argumentum ad populum, I have the impression that most computer scientists, at least, do not find Searle's argument at all convincing. Too many people for whom GEB was a formative book.
Saying “LLMs match understanding well enough”, is to make the same core error if we were to say “rote learning is good enough” in a conversation about understanding a subject.
The issue is that they can pass the test(s), but they dont understand the work. This is the issue with a purely utilitarian measure of output.
Well, I prefer it that way, but the spirit of "AI" seems to go in another direction, and the leadership of US government also does, so maybe times are just changing.
Humans also don't understand and are frequently faking understanding, which for many tasks is good enough. There are fundamental limits to what humans can do.
The AI of a few months ago before OpenAI's sycophancy was quite impressive, less so now which means it is being artificially stunted so more can be charged later. It means privately it is much better than what is public. I can't say it "understands," but I can say it outclasses many many humans. There are already numbers of tasks based around understanding where I would already choose an LLM over a human.
It's worth looking at bloom's taxonomy (https://en.wikipedia.org/wiki/Bloom%27s_taxonomy): In the 2001 revised edition of Bloom's taxonomy, the levels were renamed and reordered: Remember, Understand, Apply, Analyze, Evaluate, and Create. In my opinion it is at least human competitive for everything but create.
I used to be very bearish on AI, but if you haven't had a "wow" moment when using one, then I don't think you've tried to explore what it can do or tested it's limits with your own special expertise/domain knowledge, or if you have then I'm not sure we're using the same LLMs. Then compare that experience to normal people, not your peer groups. Compare an LLM to people into astrology, crystal healing, or homeopathy and ask which has more "understanding."
The claim was LLMs understand things.
The counter was, nope, they don't. They can fake it well though.
Your argument now is, well humans also often fake it. Kinda implying that it means it's ok to claim that LLMs have understanding?
They may outclass people in a bunch of things. That's great! My pocket calculator 20 years also did, and it's also great. Neither understands what they are doing though.
This is what you wrote:
> LLMs don't understand.
That's it. An assertion of opinion with nothing else included. I understand it sucks when people feel otherwise, but that's just kinda how this goes. And before you bring up how there were more sentences in your comment, I'd say they are squarely irrelevant, but sure, let's review those too:
> It's mind-boggling to me that large parts of the tech industry think that.
This is just a personal reporting of your own feelings. Zero argumentational value.
> Don't ascribe to them what they don't have.
A call for action, combined with the same assertion of opinion as before, just rehashed. Again, zero argumentational value.
> They are fantastic at faking understanding.
Opinion, loaded with the previous assertion of opinion. No value add.
> Don't get me wrong, for many tasks, that's good enough.
More opinion. Still no arguments or verifiable facts presented or referenced. Also a call for action.
> But there is a fundamental limit to what all this can do.
Opinion, and a vague one at that. Still nothing.
> Don't get fooled into believing there isn't.
Call for action + assertion of opinion again. Nope, still nothing.
It's pretty much the type of comment I wish would just get magically filtered out before it ever reached me. Zero substance, maximum emotion, and plenty of opportunities for people to misread your opinions as anything more than that.
Even within your own system of opinions, you provide zero additional clarification why you think what you think. There's literally nothing to counter, as strictly speaking you never actually ended up claiming anything. You just asserted your opinion, in its lonesome.
This is no way to discuss anything, let alone something you or others likely feel strongly about. I've had more engaging, higher quality, and generally more fruitful debates with the models you say don't understand, than anyone here so far could have possibly had with you. Please reconsider.
My favorite thing about LLMs is that they can convincingly tell me why I'm wrong or how I could think about things differently, not for ideas on the order of sentences and paragraphs, but on the order of pages.
My second favorite thing is that it is amazingly good at deconstructing manipulative language and power tactics. It is scary good at developing manipulation strategies and inferring believable processes to achieve complex goals.
And if that is so, didn't you also "just" express an opinion? Would your own contribution to the discussion pass your own test?
You might have overlooked that I provided extensive arguments all around in this thread. Please reconsider.
This is not what I said, no: I said that asserting your opinion over others' and then suddenly pretending to be in a debate is "not allowed" (read: is no way to have a proper discussion).
A mere expression of opinion would have been like this:
> [I believe] LLMs don't understand.
And sure, having to stick an explicit "I think / I believe" everywhere is annoying. But it became necessary, when all the other things you had to say continued to omit this magic phrase, and it became clearly intentionally not present, when you started talking as if you made any arguments of your own. Merely expressing your opinion is not what you did, even when reading it charitably. That's my problem.
> Would your own contribution to the discussion pass your own test?
And so yes, I believe it does.
> You might have overlooked that I provided extensive arguments all around in this thread. Please reconsider.
I did consider this. It cannot be established that the person whose comment you took a whole lot of issue with also considered those though, so why would I do so? And so, I didn't, and will not either. Should I change my mind, you'll see me in those subthreads later.
I did. You are not living up to the standard you are demanding of others (and which rarely anybody around here satisfies anyway).
Seems we are not getting anywhere. We can agree to disagree, which I'm fine with. Please refrain from personal attacks going forward, thank you.
Challenge semi-accepted [0]. Looking through my entire comment history here so far on this wonderful forum (628 comments), there seem to be 179 hits for the word "think" and 21 for the word "believe". If we're being nice and assume these are all in separate comments, that would mean up to ~32% of my comments feature these words, and then only some portion of these will actually pertain to me guarding my own opinions with them. Still, feeling pretty chuffed about it if I'm honest, I think I'm doing pretty good.
For good measure, I also checked against your comment history of 100 comments. 2 counts of "believe", 9 counts of "think". Being nice here only yields us up to 11%, and focusing on expressions of opinion would only bring this down further.
That said, I think this is pretty dumb. [1]
> I did. You are not living up to the standard you are demanding of others (and which rarely anybody around here satisfies anyway).
Please do show me the numbers you got and your methodology. (And not from the research you're going to do after reading this comment - although if it's done actually proper, I'm interested in that too.)
> Seems we are not getting anywhere.
If only you put as much effort into actually considering what I wrote as you did into stalking my comment history or coming up with new fallacies and manipulation tactics, I think we would have.
Seriously:
- not being able to put it into words how you don't think LLMs understand is perfectly normal. You could have embraced this, but instead we're on like level 4 of you doubling down.
- sharing your opinion continues to be perfectly okay. Asserting your opinion over others continues to be super not okay.
- I (or others) don't need to be free of the faults that I described in order for these things to be faults. It's normal to make mistakes. It'd also be normal to just own them, but here I am, exporting my own comment history using the HN API, because you just can't acknowledge having been wrong and not defending it, even though reading between the lines you do seem to agree with basically everything I said, and are just trying to give me a rhetorical checkmate at this point.
> Please refrain from personal attacks going forward, thank you.
Tried my best. For real; I rewrote this like 6 times.
[0] You continue to heavily engage in manipulative language and fallacies, so I feel 100% uncompelled to honor your "challenge request" proper. I explicitly brought up several other criteria, such as a sentence presenting as an opinion when read in good faith, not being utilized as an accepted shared characterization when used in other sentences, and not being referred to as arguments elsewhere. What you describe as "statements of opinion with the magic dust of "I believe"" seem to intentionally gloss over these criteria, in what I can best describe as just a plain old strawman. So naturally, the challenge was as woefully weakly accepted as I possibly could.
[1] Obviously these statistics are completely bogus, since maybe you just don't offer your opinions much. Considering your performance here so far, this is pretty hard for me to believe, but it is entirely possible and I don't care to manually pore over 100 of your comments, sorry. If they are anything like the ones in this subthread here so far, I've already had more than enough. And if I went through the trouble of automating it ironically involving an LLM, I'd be doing a whole proper job of it at that point anyways, which would go against [0].
Does that actually matter? Probably not for many everyday tasks...
(Besides, we know what LLMs do, and none of those things indicate understanding. Just statistics.)
You can explain this to an LLM
The LLM can then play the game following the rules
How can you say it hasn't understood the game?
Claiming anything else requires a proof.
https://arxiv.org/abs/2206.07682
https://towardsdatascience.com/enhanced-large-language-model...
https://arxiv.org/abs/2308.00304
(and if MoRA is moving the goal posts, fine: RL/RT)
That statement reveals deep deficiencies in your understanding of biological neural networks. "electrical activity" is very different from "pre-programming". Synapses fire all the time, no matter if meaningfully pre-programmed or not. In fact, electrical activity decreases over time in a human brain. So, if anything, programming over time reduces electrical activity (though there is no established causal link).
> I sometimes think this debate happens because humans don't want to admit we're nothing more than LLMs programmed by nature and nurture, human seem to want to be especially special.
It's not specific to humans. But indeed, we don't fully understand how brains of humans, apes, pigs, cats and other animals really work. We have some idea of synapses, but there is still a lot unclear. It's like thinking just because an internal combistion engine is made of atoms, and we mostly know how atom physics and chemistry work, that any body with this basic knowledge of atom physics can understand and even build an ICE. Good luck trying. It's similar with a brain. Yes, synapses play a role. But that doesn't mean a brain is "nothing more than an LLM".
Humans arrive out of the VJJ with innate neural architectures to be filled and developed - not literal blank slates, there is an OS. The electrical activity during development is literally the biological process that creates our "base programming." LLMs have architectural inductive biases (attention mechanisms, etc.), human brains have evolved architectural biases established through fetal development. We're both "pre-programmed" systems, just through different mechanisms.
Your response about "electrical activity decreases over time" is irrelevant - you weren't talking about adult brain activity, you were talking about the developmental process that creates our initial neural architecture.
tbh: I can't tell if you're engaging in good faith or not.
An LLM can pass many tests, it is indistinguishable from someone who understands the subject.
Indistinguishable does not imply that the processes followed match what a human is doing when it understands a subject.
I use this when I think of humans learning - humans learn the most when they are playing. They try new things, explore ideas and build a mental model of what they are playing with.
To understand something, is to have a mental model of that thing in ones head.
LLMs have models of symbol frequency, and with their compute, are able to pass most tests, simply because they are able to produce chains of symbols that build on each other.
However, similar to rote learning, they are able to pass tests. Not understand. The war is over the utilitarian point “LLMs are capable of passing most tests”, and the factual point “LLMs dont actually understand anything”.
This articulation of the utilitarian point is better than the lazier version which says “LLMs understand”, and this ends up anthropomorphizing a tool, and creating incorrect intuitions of how LLMs work, amongst other citizens and users.
It’s incredibly easy to get LLMs to do a lot of stuff that seems convincing.
They are literally trained for plausibility.
For instance, you might have an SEO expert on the team, but that alone won't guarantee top search engine rankings. There are countless SEO professionals and tools (human or AI-powered), and even having the best one doesn't eliminate the underlying challenge: business competition. LLMs, like any other tool, don’t solve that fundamental problem.
This sounds like those guys in social media that one up each other with their bed times and end up saying they wake up every day at 2am to meditate and work out
And other companies have existed for hundreds of years and had thousands of people work for them and never even made $100M.
This is epic work. Would love to see more of it but I guess you're gonna take it the startup route since you have connections. Best of luck.
I was just building a tool people can use to do the business side of product and technology. I wanted it to do basic tasks i KNOW a business needs at all scales, not pretend to do whatever "AI" thinks an employee does.