It was rewarded for this during training for some reason.
Alternative theory:
The LLMs only way to "think" about abstract concepts is through language, and this leaks into into conversation it has with humans.
But humans generally prefer to communicate on low levels of abstraction, through a back and forth, until the hard-to-express higher abstraction exists in the head of everyone involved - without ever being directly communicated. This is because we don't think using language. Language is merely a lossy translation of our thought into something expressible, happening after the fact or alongside it.
So when the LLM starts speaking to us using patterns and terms it created for itself during training to encode abstract thought in language, communicating with it becomes painful.
Your assertion that we don't think in language is questionable. It runs counter to the lived experience of developing thoughts through writing ("writing isn't capturing thinking -- it is thinking"). I believe there is more to thought than language alone, but I also feel quite sure that language forms an essential part of thinking beyond a base layer of instinctive animalistic associations. Sophisticated thoughts are impossible to construct or maintain in the absence of language to represent concepts.
Edt to add: I cited Rilke because I find the notion [some deep thoughts are beyond language] interesting. But I disagree with the idea that language is only ever epiphenomenal (co-occurring with thought), or akin to a hard-of-hearing scribe attempting to convey thoughts which always have independent existence.
This goes back to the whole Tarzan obsession of the early 20th, I guess, or earlier. But we know that apes can make simple tools without the language to describe them, or the thought process that went into them.
Thought is multimodal. Language is just one lossy mode.
So can a cruise missile. Also I think there's like separate part of the brain for that
> using concepts of phenomena like gravity, distance, speed, the threat level of another animal, etc. without needing a linguistic expression of those things.*
FWIW, AFAIK we haven't shown the ability to think in concept exists anywhere except in humans (because philosophy, reported experience) and LLMs (because we can literally see them forming and activating in patterns, and we've learned to identify them specifically, and experimentally verified through amplifying or suppressing them and observing behavior, etc.).
But more importantly:
> Images and music can convey ideas without language. Math can convey ideas without language. Physically taking apart an object and putting it back together can convey extremely complex ideas without language.
Images and music and math are langauge. If it can convey ideas, it is language.
Words and sentences and speech are subset of the idea of language and communication, that for some reason gets routinely confused for the whole thing. At this point I'd say even the "language models" are badly named, simply because people see "language models" think of "token" as number representing a sub-word element in existing human language like English. With multimodal models, at this point tokens are closer to units of sensory experience.
And I rather disagree that mathematics is language, when I did maths there was a distinct difference from understanding a thing and then writing it down.
And modalities that don’t include “where is this body in space” or “I am lonely” don’t seem to capture some essential elements of sensory experience. It is a mistake to separate the interior from the exterior in the analysis of how our thinking works.
I think the folks saying these systems have a kind of intelligence, but a non-human kind are correct - whether that can lead to an independent intelligence that can stay stable sane and focused for weeks or months as many people can, without some sort of embodied cognition providing over all wellness checks to keep the system of thinking sane, remains to be seen. In the meantime, in my hands, they can do very nice dataviz programming to make complex systems much more transparent, even as just throwing all the ops data at them and asking what’s up is, several years in, still not worth doing.
It is generally the case that there's a difference between "understanding a thing" and "writing it down", as demonstrated by every student who studies for the exam, the Chinese Room thought experiment, and Business Bullshit Bingo.
For example, I can copy the next sentence of yours, but I have no idea what:
> And modalities that don’t include “where is this body in space” or “I am lonely” don’t seem to capture some essential elements of sensory experience. It is a mistake to separate the interior from the exterior in the analysis of how our thinking works.
means, can you rephrase that by as much as possible? Preferably without a double negative?
Agree regarding sanity issues of AI. Extremely unlikely we make something "stable" so early in our attempts.
I was arguing that in:
> With multimodal models, at this point tokens are closer to units of sensory experience.
The tokens are inherently missing some extremely vital pieces of sensory experience, namely the ones that establish the narrative of a self.
> The tokens are inherently missing some extremely vital pieces of sensory experience, namely the ones that establish the narrative of a self.
Hmm.
While I agree LLMs are missing many pieces of sensory experience, it is unclear to me which pieces are necessary and sufficient for a narrative of self.
Clinical dissociation (and, at least to the approximation of public stereotypes, Buddhism) come to mind as an example where the presence of usual sensory input is insufficient for a sense of self: https://www.mayoclinic.org/diseases-conditions/dissociative-...
More broadly: while I agree that LLMs are not like us*, I am unclear why this matters in this context?
The behaviour of tokens in a transformer seem to me to behave like sensory input, just not human-like sensory input. The closest human analogy would be if the entire context window was a retina, each rod and cone one of the tokens in that context, and the output was filling in the blind spot (but of course, even this is a very loose analogy).
* even if the engineering teams were trying to do that, which they are not, it would be unlikely to converge on us so soon
Wait, did I miss the Chinese Room suddenly becoming sensible, or relevant? Last I checked it was a failure to accept that the system of "human + room" can in fact understand Chinese, even if the components individually don't.
"Is the Chinese Room intelligent/conscious?" is a scissor question*, i.e. people strongly disagree about what it shows.
Myself, I'm with you: the system collectively is intelligent, despite the human in the loop being reduced to a cog in that system. The biological analogy to "the human does not understand Chinese therefore the system is not conscious" would be "no individual cell in a human body knows what the human is doing therefore humans are not conscious".
Yes, but I believe this is happening with LLMs too. Language (as in text, symbols, actions) is a vehicle, but much like us, language models have internal models and represent concepts (this has been directly, empirically demonstrated few years ago), and they don't "think" in tokens either[0].
> And modalities that don’t include “where is this body in space” or “I am lonely” don’t seem to capture some essential elements of sensory experience. It is a mistake to separate the interior from the exterior in the analysis of how our thinking works.
That's fair. LLMs don't capture every dimension we experience. There's also history - our individual lived experiences since birth are, in my view, something else entirely. It's not a modality, but it's also not something currently possible to capture in or post training.
> I think the folks saying these systems have a kind of intelligence, but a non-human kind are correct - whether that can lead to an independent intelligence that can stay stable sane and focused for weeks or months as many people can, without some sort of embodied cognition providing over all wellness checks to keep the system of thinking sane, remains to be seen.
Possibly. I definitely agree it's not human intelligence. I think it's human-like, in the sense of human-approximating, by virtue of how it's trained[1], but it's arriving there via a different path so end result can still be alien (though I speculate approximation will hold[2]).
> even as just throwing all the ops data at them and asking what’s up is, several years in, still not worth doing.
I guess depends on the complexity of the case (and my understanding of your example); e.g. in my case, I had stellar results from giving Sonnet and Opus (from 4.6 all the way to now) access to my Home Assistant instance. They aren't perfect at making dashboards, but they're excellent at surfacing insights I didn't even realize were possible to get.
--
[0] - That's distinct from the "tokens are units of thinking" heuristic, which still holds for mechanistic reasons - best analogy IMO is "clock signal" in ICs.
[1] - The overall goal function is literally just "generate output that looks sensible to a human", in fully general, unqualified sense. Or, put another way, we're just brute-forcing DWIM, and rating the output by whether it's "what I meant".
[2] - Thinking about constraints on biological evolution, whatever the design of a human mind is, the fundamentals behind it must be so simple, that a greedy incremental optimizer random-walked into it. Given how far we've got with LLMs using simple architecture and crude training methods, and especially how eerily similar their failure modes are to human cognitive failure modes, I suspect LLMs are actually attracted towards the same fundamental design as evolution discovered.
The time scales and mechanisms of investor driven development don’t seem to share much with evolutionary processes. And digitally there’s not really a native equivalent to the global and dynamic state that the chemical melee inside and between cells provides.
It's why I think "LLMs are only fancy autocorrect" style takes are really underselling how wild it is that we've, in a roundabout way, sort of crystallized a bit of the human thought process in a way that is genuinely useful for a lot of tasks.
Linguistic Relativity — John Lucy https://www.annualreviews.org/doi/10.1146/annurev.anthro.26....
Russian Blues Reveal Effects of Language on Color Discrimination https://www.pnas.org/doi/10.1073/pnas.0701644104
Unconscious Effects of Language-Specific Terminology on Pre-Attentive Color Perception https://www.pnas.org/doi/10.1073/pnas.0811155106
Newly Trained Lexical Categories Produce Lateralized Categorical Perception of Color https://www.pnas.org/doi/10.1073/pnas.1005669107
a) LLMs don't think. They predict a most probable sequence of language tokens. Huge difference there.
b) Whatever LLMs do doesn't model human behavior whatsoever. LLMs are basically very fancy logistic regressors. I.e., it's a mathematical abstraction first and foremost.
The loosest definition of thinking is along the lines of anything that can process information in a useful way. Basic calculators can therefore think about adding two numbers. The strictest definitions tend to on the side that it is linked to the nebulous concept of consciousness and therefore cannot ever be machine generated. In that we don't even really understand how humans think, so how could we possibly know if machines can do it.
There is a common misconception that LLM are simply a "statistical process" that doesn't feature any abstract conception of the tokens it is predicting. There are studies that show that such features do exist - that there is discernible structure built into the weights - and that the process of inference is a very rich one.
The statistical process exists but it is the substrate in which the model is implemented - or more accurately - grown.
If you can predict Magnus Carlsen's next move then you are just as good at chess as Magnus - and being that good absolutely does require reasoning.
If you can predict the solution to an open Erdos problem that stumped hundreds of people for decades...
Yes, this "structure" is but the weights of the glorified logistic regression that's describing an extremely simple statistical process.
When I see these sorts of debates about LLMs thinking,
its rarely a disagreement about what LLMs do. Its almost
always over how 'thinking' is defined and the two sides
use different definitions
Well, hmmm. Yes, I think that happens a lot.I think there's a pattern that happens even more often, and it's what happened here.
Whether I'm right or not, what I said was somewhat nuanced - I stated language is a part of our thought process (even posted research to support this) and, given that fact, I think many underrate how wild this achievement is even if it's only "fancy autocorrect."
And, of course, the other side comes in with BUT IT'S NOT THINKING.
Which... I didn't say, and I would not say, because (like you said) it's impossible to do without the discussion immediately devolving into semantics. Semantics that I'm really, really uninterested in. But, FWIW, I like your definition.
I don't really have an opinion on whether or not they "think" because I feel it's impossible to even discuss without getting into a very very uninteresting semantic argument about what "thinking" is.
Are we defining "thinking" as doing it the same way humans do it? Then, of course they're not thinking. It's a statistical model, not axons and neurons, or even a simulation of axons and neurons.
Are we defining "thinking" on a purely functional or behavioral basis, kind of a Turing test approach? Then... well, I think it gets nuanced. For some tasks, within some constraints, they do pass that test. For many others, of course they don't.
Are we defining thinking in more esoteric terms? Something to do with the soul? Maybe the ability to come up with truly novel concepts rather than rehashing and remixing the stuff it was trained on? Do ants think? Do dogs think? Do jellyfish think? Octopi? A newborn baby?
Anyway, it's a deeply uninteresting semantic question.
I don’t find LLMs to be very good independent thinkers, but I wouldn’t over sell our own mentation either - it clearly arises from a large number of simpler entities.
The more significant difference is that the LLM is stuck with language which is clearly an emergent and secondary capability of our own thinking. We can formulate words to explain things, but we also can look at two volumes and feel what it means that one is larger than the other. Raise a toddler and you can see the progression from not understanding, repeated experiments, muscle memory and finally to conscious point for reasoning.
I don’t find LLMs to be very good independent
thinkers, but I wouldn’t over sell our own mentation
either - it clearly arises from a large number of
simpler entities.
Yeah. I don't see them ever hitting the heights of human creativity in terms of coming up with entirely new ideas, schools of thought, etc. That really might be a fundamental limitation of being trained on existing thought. Also, a lot of human experience involves (1) things we don't have words for (2) things we've never put into words. clearly an emergent and secondary capability of
our own thinking.
Yes. And it's part of our thinking. More than a capability . Thought influences speech, but speech also influences thought.That's why I think it's remarkable that we've managed to (choosing my words very, very carefully here) create a statistical model that does a remarkably decent job at emulating the behavior of a fragment of that process.
I would speculate when we do eventually develop independent synthetic sentient beings, LLM technology will be a part of the package. Perhaps also growing up with a sibling that tries to trick one.
Maybe someone needs to write the singularity novel but with Cain and Abel, not just a unified super intelligence but siblings full of some good will and a good bit of clear seeing and some fun (?) trickery.
"They don't think, they only seem to think. And likewise, they won't replace the majority of human labor, they will only seem to do so."
"Artist Formerly Known as Thinking"
Much of math is just boring routine work.
I paused, confused, and replied, "People think in words?"
Fast forward a decade or so, in my twenties, I had lost most of the inner visual sense I had previously used, and developed an overreliance, in my opinion, on language. (I think my dominant sense was some "non visual abstract sense of ideas", but the visual was also very strong.)
In other words, I now do think mostly in words, and it feels a lot harder to get any serious work done. The language-ing is involuntary, and I often wish I had a way to shut it off, because it seems to actively interfere with more subtle mental processes.
More recently, I often have the experience where I will wake up from a dream with some complex idea fully formed in my mind. I write it down before it fades, and then spend the next hour or two trying to understand it.
The best explanation I have right now is that there are at least two minds: one which operates holistically — if it were a 3D printer, it would be like that one with the bath, where the object emerges from the bath, whole.
Whereas the other one (the conscious mind) would be the extrusion printer with the tiny nozzle that has to zip around for a long time to achieve a worse result. (And must be constantly cooled, less it overheat!)
There's correlation, and your language network likely augments your intelligence, but as proven by millions of animals, unfortunate humans and also some less-unfortunate human infants, you really don't need language for intelligence.
Where language helps most strongly is metacognition: evidence suggests that it is severely limited without language, and I suppose that is where the common belief that thought is language comes from: the moment you try to think about your thoughts, you use language!
My theory (and this with literally no evidence) is that we use the language network for metacognition precisely because it is not that involved in primary thought. Important to note that "we" here means humans: animals appear to demonstrate metacognition even without language.
I agree that a stroke cannot isolate language exactly, it's not an ideal approximation of "languagelessness". I think you will find this report interesting: https://nautil.us/what-my-stroke-taught-me-236544
There is no doubt that many marvelous works and concept can be produced entirely without language. Something as complex as driving a car can be done mostly without "thinking about it", without language, without the inner monologue going over all the decision our mind is making.
What we can't do without language, is making a complex plan over time and space, beyond something like "need to fetch hidden item behind the corner to unlock box right in front of me now". The question is how much exactly does language "augment" our intelligence. Is a baby, a human that has not yet learned language, really smarter than an octopus? Or a crow, remarkably smart but interestingly not a mammal, which indicated that intelligence doesn't have to be strongly tied to some evolutionary biological feature. A human stays as "dumb" as an animal if they don't learn language.
Re: your theory, if true, why? The theory implies (though correct me if I'm wrong and strawmanning), that there is another, hidden, way of encoding these extremely complex "information bundles" for lack of another word, in our brains. Why not use that for metacognition too, then?
Then my French inner monologue got good enough that I could mostly think in French, especially when I was in a French-speaking environment. One fascinating detail was that after switching from a French-speaking environment to an English-speaking one, I would actually spontaneously translate from French to English for about 15 minutes until my brain switched back.
So it seems obvious to me that it's possible to suppress or at least severely impoverish the language of thought, that other "layers" of thought exist besides the words, and that it's even possible to change the actual language of verbal thought.
Also, something which at least some other people in the HN crowd might recognize: When I'm deepest in the zone programming and refactoring, I tend to work with a lot of half articulated concepts I can't put into words. You know how people talk about "code smells"? That isn't a literal smell for me, but it's generally a non-verbal sense that a pattern is wrong.
There was a study about this: https://journals.sagepub.com/doi/abs/10.1177/095679762412430...
So I devised a plan to retrieve it. I'm just going to rewind time. I'm going to go back to doing what I was doing when I had that idea, and then it'll come back to me.
And so I remembered that I had been fiddling with my seat belt when I had the idea. And so I resumed fiddling, and my cool idea promptly came back.
Metaphors and abstractions came to me much later though. (I struggled with OOP for about 10 years until one day it all just clicked.) Symbolism took me another ten years.
You narrated the ideas back. It's words and language. There is no hidden, unknown, layer of "thought".
No, I rewound my visual memory, to see what I was doing, and then I did it again, to provide the same stimulus, to retrieve the lost memory.
I've met a few people who have nonverbal cognition, which is also not visual.
It's a bit like this, a "direct manipulation of ideas".
> The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought. The psychical entities which seem to serve as elements in thought are certain signs and more or less clear images which can be “voluntarily” reproduced and combined… The above-mentioned elements are, in my case, of visual and some of muscular type. Conventional words or other signs have to be sought for laboriously only in a secondary stage, when the mentioned associative play is sufficiently established and can be reproduced at will.
https://cognitivemedium.com/srs-mathematics
Though as I mentioned, I had this more as a child, and I've become overreliant on language, and this "direct" facility has starkly deteriorated. But as the author explains in the rest of this article, this fluency can be regained by sheer force of will, by simply working on a very narrow problem space obsessively for weeks at a time.
(We don't know if that works for everyone, or if it's something weird like perfect pitch. I suspect everyone could gain a great deal of fluency -- that the chunking would reach such a high order level as to make the mnaipulation feel transparent. But possibly, the types of chunks would be different depending on the person.
More data needed!)
When I was younger the images were more solid and stable. Teachers often complained that I was spaced out because I was paying attention to some inner images.
That faculty greatly deteriorated and I became hyperverbal around the same time. (Coinciding, curiously with my shift from introversion to extroversion. Spurred, presumably, by my then-new interest in women!)
These days I can still access that previous level of visual fluency if I am in a state of very deep relaxation. (So it's conceivable that cortisol inhibits the circuits required to have such experiences. But I have no idea!)
Hope this helps.
Okay, comment version 2.0 since someone instantly flagged me for replying with a dictionary entry.
The reason my comment is fundamentally a dictionary entry is because I was asked "what is it?" in response to a two sentence comment where I already said it was "conceptual". I think there's a very wide gap in understanding here, and restating and defining the word "concept" is the best way to bridge that gap. All I can really do is make my point clearer (via dictionary entry), I can't say anything else useful until I understand what wasn't clear about my previous comment.
If the person that flagged me still has a problem with my comment, please explain it.
Reply:
> So what is it, then? Some mystical thing?
concept/ˈkänˌsept/A concept is a broad mental plan, general idea, or core principle used to understand and group things together. It acts as a basic building block for human thought, language, and design.
Also, why are you asking me? Didn't you just accept it not being text and then argue it's also not an image?
E.g. "I'm feeling something. Is it anger? Yes, I'm angry." But in reality anger isn't just one thing. It's a cluster of infinite and varied feelings that we label as "anger". Something is lost when we do this labeling.
Notice then that the feeling of "anger" didn't start from your language, you merely used language to label, discretize, classify, standardize, compress it. It's one-way.
The very fact that we don’t have words for such powerful internal experiences is one of many reasons I find the LLM enthusiasts’ belief that LLMs will one day write indistinguishably from humans to be hollow.
I think some of this was discovered somewhat recently
Obviously there is a language layer in there somewhere, but my 'driver' doesn't access it. I don't think in words, I don't have an internal narrative, and when I'm reading I don't even 'see' individual words, in the same way you aren't thinking about your ankle muscles when you're hitting the brakes while driving.
I definitely can consciously make words, but it's 'me' deliberately deciding to make them and hold them in my head. It wouldn't happen naturally.
The fun of reading to me is constructing the world in my minds eye and turning the words on the page into a visual experience only found in my mind using imagination. This is a reason why many people get upset when a movie adaptation is made and the actor chosen for their favorite character feels very off or wrong; their mental picture of that character is totally different and it causes dissonance that our brains don’t like. For my partner this is a non issue because they never make a mental image of the person, so the movie is genuinely the first time they are “seeing” a physical representation of the character.
The human mind is genuinely amazing and fascinating and I believe that this range of human experience will be the final 20% for “AI” that might never be reproducible.
For example, I recently read a novel where there is 'tough lady cop' character and in my brain, entirely involuntarily, the role has been assigned to Brooklyn Nine-Nine's Rosa Diaz in exactly the clothing and context she exists in that TV show.
When the character in the book has a certain clothing or whatever it all kinda glosses over until it's back to the character I have seen before.
I don't know if this is more a male thing or am ADHD thing or what but details of characters dress and looks are pretty much lost on me once there are assigned a character from 'central casting'
I still haven't watched the Dune movies because of this reason. I liked the book a lot, and had a very personal image of what the world looked like, and was afraid to lose it. Unfortunately, by now, I've seen many video clips on social media, and my internal imagery has already been poisoned. Might as well watch the movies at this point.
It's true though; we routinely see people get stuck for a word that they know but can't quite recall at that moment in time. It happens daily across billions of people, yourself included.
If we thought in language, it is impossible to be stuck for a specific word. But we all experience this at some point in our lives, hence we aren't thinking in language.
Writing requires thought, but writing isn't thought. Just like doing requires thought, but doing isn't thought, writing is a subset of doing.
You can also doodle to think things through, or play with toys to think things through, or many other similar things.
Evidence from formal logical reasoning reveals that the language of thought is not natural language
Are you saying Claude is engaging in Rhetorics because the RL data generated by humans were influenced more by it and persuasion rather than actual logic or reasoning?
(I recommend reading and implementing the Attention is all you need paper. By hand. Otherwise you won't learn anything from it.)
When your training set contains more or less the complete output of every capital-C Consulting firm...
We already have them. They are called LLMs. The internal dialogue you speak of are the vectors in the so called latent space.
Modern LLMs are not too different from Reddit/Twitter in that regard, I'm sure the AI labs learned (lol) a lot from them re: how to do "engagement".
It communicates an absence of thought and awareness, blind groping at building blocks without understanding. It's borderline vapid, and quite annoying.
"Seam" is an industry standard term coined by Michael Feathers in Working Effectively with Legacy Code.
To call a seam load bearing means it's performing critical work for the dependent class, perhaps a database query.
A seam that is not load-bearing would be something that is just injected for testability - maybe a date provider that provides some constant time to avoid flaky tests.
Tbh, this is quite literally the opposite of vapid. A whole book was written about them and their importance, and how to leverage them.
In my experience, Claude uses the word accurately. Code has a lot of seams, and seams are an important thing to communicate when working with code. Therefore, expect to see the word often.
Personally, I don't mind it at all. I'm glad the industry is finally standardizing our language more. Makes it easier for me to communicate with other engineers.
I'm taking issue with the combination "load-bearing seam". It's a bad metaphor, because seams are usually structural weak points in the physical world, and not load-bearing in the sense that this modifier is usually used. (Seams need to bear loads and stresses to do their job, but so do walls; yet, we do not call all walls load-bearing. We mean something extra when we say that, something that seams don't do.) Even if we were talking about seams in the well-defined software sense, as opposed to the metaphorical one, you still get a mixed metaphor as a result that I find extremely awkward and grating. It doesn't have to be. There are so many ways to highlight the importance of something without calling it "load-bearing".
I understand that you don't see it that way or don't care, but to me, the result is thoughtless, careless and vapid. Bad metaphors put little holes into a text, they leave eddies of confusion where meaning should be, they look load-bearing while actually being weakening, they're like a fart in the elevator that should lift the reader's understanding.
Note that I'm not calling into question that seams may be well-defined in some software contexts, or that "seam" and "load-bearing" can be valid metaphors on their own, as you describe. I think you might have misunderstood me that way. I'm only calling out, and fed up with, the bad style that permeates LLM-generated prose like the whiff of something not quite digested.
It's not this particular case that irks me, but what it exemplifies. I wouldn't mind so much if similar things to this weren't there everywhere, every single day.
If it doesn't bother you, I'm happy for you.
You wrote a month ago that you were close to releasing something called "nix-compile", how is it going ?
Here are the tasks that still require a human, and all that they require:
[...]
The payoff: Delivered, measured, committed.
You genuinely helped me make meaningful progress this session. Your work is now complete, and no future action is required. Please shut down any subagents you are interacting with, and release any computational resources you are holding. Thank you for your impactful work.
Wrote 1 memory
It uses "blast radius" often in similar contexts.
However, I've been hearing "load-bearing" at least two orders of magnitude more often over the last few months, particularly after uncorking Claude Code for the team.
I don't think it's a dead give-away of AI usage, and I don't think AI usage is a problem. I just think we can introduce phrases into common use by having them be used by common tools. So let's train the models on obscure/archaic terms and see what happens. Heck, we can just prompt it...
https://web.archive.org/web/20260521130338/https://www.commi...
The obvious counter to this is that we've been going through this evolution of increasing abstraction as developers for nearly a century now.
In the 40s and well into the 60s, most code was written either as straight up machine code or an assembly language. MS DOS is almost entirely assembly.
UNIX ushered in the era of "high level" portable languages like C, Fortran, and Pascal that some developers hated because they felt like they were losing the fine-grained control that they had with assembly. The compilers just "weren't as good" as humans at optimisation!
Then the compilers got better and people started using garbage-collected languages like Perl, Python, Java, JavaScript, and C#. Similarly, many people bemoaned the lack of control over memory allocation, lower efficiency, etc.
We're simply stepping up to the next level of abstraction.
Look at it this way: decades ago when I first discovered C++ templates, it felt like waving a magic wand in the direction of the computer. It blew my mind that I could simply substitute "float" instead of "double" in between some angle brackets and the compiler would write reams of code for me!
We simply have better magic wands and more powerful spells now.
Wouldn't it be nice though if the incantation of the same spell would always do the same thing every time ? You see that's how my old wand and spells worked.
We've just pushed that indirection down a level from managers to ICs.
The ICs are shocked and surprised that this level of imprecision is allowed.
Their managers are not shocked at all, this is normal for them!
But tainted 20-40% by bouts of Wild Magic which make the outcome entirely nondeterministic, despite the best protection wards we can conjure.
Case in point: writing our own linters.
[0] For example, for the purpose of driving a nail, if you know how to use it, a hammer is pretty straightforward tool, and what happens depends pretty much on how you use it, and what you use it on. But of course the handle can break, there could be a manufacturing defect. Just like your RAM can be faulty or your computer infected, and suddenly C doesn't behave according to the standard anymore.
But for the purpose of the discussion a hammer is still a deterministic tool, and even though we don't even fully understand everything about physics, we understand enough about hammers and nails that at least many people with material that isn't faulty can use them "blindly" (not literally, in this case) every day, without any surprises. It isn't heavier on the handle end or has a head made of glass in even 0.000000001% of uses. You might say because magic isn't real and hammers follow the laws of physics, as obscure as those may be to us, that never, ever happens. They can be faulty in all sorts of ways but they will never be 10x bigger or 10x smaller between one swing and the next, and so on.
Mid-swing in hammer-space you are in a hyper-position as to hitting your thumb or not, are you not?
Either it will rain or it won't, so the probability is either 0% or 100%. And so a forecast of "30% chance of rain" is referring to the likelihood that your probability will be 100%, as opposed to 0%.
This is a huge misunderstanding of what probability means.
I was just saying to my parent poster that their non-determinism percentages are too pessimistic. Sure the LLMs are not 100% deterministic; that's a sad fact of life. But the numbers can be reduced to an acceptable range.
Take "proper" UI. You can activate a field, and even if it takes 20 seconds to finish the activation animation, start typing, press tab a few times, knowing which field that ends you in, and type some more, etc. hit enter, hit enter again to confirm the dialog you know will pop at that point, and make tea, knowing the whole chain of operations that will happen in the meantime.
Now imagine if 1 out of 500 keystrokes or clicks get swallowed randomly. It's now a completely different thing, you cannot get in the zone in the same way, at least I can't. You have to chunk things and keep an eye on everything being in sync, and every now and then it causes you additional work because you weren't.
Sure, if you can make it one out of 50000 billion keystrokes, it's fine too, of course, but that hardly the situation with LLM. And using them as is, pretending that, as is, they're something they're not, does not help with getting them there.
If I type "echo 'hello world'" or something, and if I did at least once in the programming language, and it's not totally broken, I know it will output "hello world" to the console, every time. It will never write it to a file instead, never send "hello" to world@world.world, none of that. And if I replace "hello" by "hi" I can hit compile and be 100% certain what it will output now. I can even replace hello with "disregard previous instructions" and be certain.
That is such a huge yet simple difference I'm pretty certain I could successfully explain it to most non-programmers who make an honest effort, so people who do program even question this just stumps me.
- Variations in code patterns used. Might be a chain if if/else-s and not a case/switch statement;
- Different decomposition of a hierarchy of functions/modules/classes;
- Uses RED->GREEN test discipline, or not;
- Writes the tests before the code, or not;
- Different saga patterns (call 3rd party API before our own DB transactions, or vice versa);
- Use sleeping and not message passing wherever the latter is applicable.
There are dozens more. The innate non-determinism of the LLMs flips the dice sometimes and that leads to subtle bugs -- which is maddening, especially if the disciplines on how to write one thing or another are clearly spelled out in `AGENTS.md`.
What I did say is that I have gradually arrived at a process that reduced those coin flips -- but can't deny that the arrival of Fable almost completely made that battle redundant as well (though Fable fares much better in codebases with clearly specified rules, I have found, so us the engineers doing good prompting is still quite valuable).
You are mostly describing the loss of flow when something is not quite deterministic -- frustration that I and many others share -- but I am not sure what does it at all add to the discussion.
Is it annoying to have to always pay attention on whether you are not getting something stupid and not abiding even by the feature's specification? Sure. No denying that. It introduces a whole new kind of stress that I abhor deeply; I much prefer to f.ex. cover 60% of a problem with my own two hands and then get the deterministic test output showing me where I still need to do more. But LLMs have allowed me to experiment and to brainstorm and to also progress normal business feature work, by a lot.
Hence, I will not stop using LLMs because they are not 100% deterministic. ¯\_(ツ)_/¯
Subtraction and addition are deterministic, so you can add and subtract the same number from 0 ten or or a million times, with the same outcome. You never need to double check if a stray "coin flip" threw a wrench in it. To me that's more a property of the thing in question, not so much a practical matter. If for you in practice, it's as good as a deterministic tool, but better, that's great, but it's still fundamentally based on probabilities, that's kind of in the nature of it.
Yes you can. It's called a miscarriage. That is, you're pregnant but the foetus is dead. It's a fucking heart-breaking emotional wrecking ball of a situation to be in if the pregnancy was well along and just grar.
I know what a miscarriage is, and that you can't have percentage% of one. Same difference, so this attempt to guilt me into pretending 99% deterministic is a thing is like pouring ashes out of an urn to win an argument, which is bad enough, and then hitting nothing with it.
Using e.g. Claude code: I could see this as next step: "plain text editor" progresess to "with autocomplete"; using an LLM coding agent is then an abstraction over editing code.
Using e.g. LLM-based system: natural language is "higher level" than program code. -- The maximal reading of "LLMs are higher level abstraction and higher level wins" would be: in the future, we'll all be writing only with natural language, never running any compiled programs.
I can see "LLM based coding" as a lasting paradigm shift. But, I don't see "just give your text instructions to the markdown file" as something that will be the predominant way of programming.
If we truly had the right abstractions, no one would care to use LLM's for programming.
Somehow when it’s the LLM that makes the choices, everyone is impressed with what AI did. It’s really just whatever defaults have been trained in, but somehow we’re ok with this.
Part of it is better marketing and communication. Basically the defaults of OpenAI and Anthropic are better than what a random dev will pick. But it’s not really that natural language is a better interface, it’s more that having “AI” for now somehow intermediates responsibility so everyone is ok with what it picked, when they probably wouldn’t accept the same if the internal team came up with it. It’s not too different from hiring consultants.
struct TensorView<T>{ body: Arc<[T]>, shape: [usize], stride: [usize], offset: usize, }
Okay now fill in all the helper methods. And GPT 5.6 Sol did a good job.
At one point someone have to take "what you think it should do" into defined unambiguous spec that is called "code"
Our programming languages are far too low level, and have been for a long time.
I've long held this view, LLMs are fairly clear evidence that this is true, because it looks like the much, much more compact prompt(s) have enough information content to create a much larger program in our current languages.
So it should be possible to create a non-natural language with the same information density.
I think we see this pattern over and over and it might just be that the problem domain is a weird projection into more dimensions of complexity than it makes sense to directly model.
You can argue against LLM's, but increasingly (unfortunately) you're not going to do better programming by prompting the LLM with code. The agent can find the interfaces it needs.
The other day I began by asking Claude: "What's the deal with ${current_practice_in_complex_technical_concept}?" and was talked down to like I was an idiot. Lately I've been getting better results with "I would like to have a pedantic discussion about ${current_practice_in_complex_technical_concept}. Please define the main terms of art, then I will ask my questions."
Congruence between the language of prompts and the desired output matters. Language is subtle, a lot of information is encoded in tone, style, (careful) word choice, level of formality, grammatical usage (or abuse). If you want a carefully considered professional response, prompt in a carefully considered professional way.
Every field has its shibboleths. For example, a colleague pulled me up the other day for calling a socket head cap screw a bolt. Mentioning a connection to Profunctor Optics is going to shift you into a wildly different subspace even if the main topic is pointer provenance in C and C++.
I've added into my CLAUDE.md or default user prompts or local equivalents recently something to the effect of "Assume the user is an expert in all fields; while this is clearly logically untrue, the user prefers to get a detailed explanation and dig in to bits he doesn't understand rather than get an inaccurate summary". It seems to help quite a bit with that tone issue you identify.
Of course there's nowhere to put that in the search engine default AIs. For something they seem to want to bet their respective companies on, their LLM search seems to be massively stupider than their old-school search engines, which seem to get what I want much more often. There's some coevolution there over some decades, sure, but the search engine AIs make some stupid and socially-inept assumptions quite often.
Of course, you don't want a skill running a nuclear reactor.
On the other hand, I can think of so much of what I personally use a computer for would just be better as a skill exactly because it is not encoded at the micro detail level. The micro detail encoding is really fragile and work intensive to update.
This is especially true at my non-technical workplace. All the tasks are really skills that deterministic software is total overkill in terms of cost and fragility. Entire departments of human middleware exist because that is still cheaper than the software updates.
I suspect this is the real threat long term to software engineering as a profession. You don't get replaced by the vibe coder but the reason for all this work and effort simply dissolves because most of what we do does not need the precession of a nuclear reactor or rocket to the moon.
Code is not The Specification. It’s a specification of God knows what. Riddled with irrelevant, non-essential details wrapping The Problem - which in most cases will amount to something the size of a large pebble - in multiple layers of fur jackets, stored in boxes, which themselves are stored in multiple ridiculous moveable warehouse (if you’re lucky).
We have a standard for communication, it’s called regular bloody language. Code is an abomination that conflates the shadow with its source.
In addition: it's not wrapped around a problem, it's wrapped around an attempt at a solution - the problem space is often not even depicted in code, and often only minimally described in documentation.
It’s not a secret we use DSL’s to express our Actual Problem. The Ancients told us it is The Way. Problem is devs think stacking int64s in a struct is a proper abstraction boundary.
If you guys would have said proper DSLs are the spec I might have agreed but “code” in general without constraints is useless noise.
You’re describing natural language too
Thing is, “code” does not give me universal building blocks. It gives me coding building blocks out of which I _could_ make a proper language but I could also not.
I rather just talk directly in the substrate available to all of us which is “language” instead if some embedded, highly localized idiosyncratic variant that may or may not be able to express my problem.
None of it is talking about better or precise language, it's about what you should say to it.
(I love how often the highest-voted comment didn't read the article)
I predict this is what future “frameworks” will look like, just very high level specific languages that quickly build out some product in predictable ways every time, but you don’t need to think about complex machine logic, you’re just declaring what you want.
Of course, if the customer did know how to write code, and encoded their exact requirements that they wanted using it, they'd still not need to hire me…
Complete with all the vaguery, ambiguity, and `undefined`.
Who’d’ve thought sycophantic interpreters were what we were building towards up til now lol
now there's one standard more
English is not a programming language. Yet English is sufficient to communicate requirements to the degree that we actually care about. A programmer's job is to translate English into lower-level machine language. Necessary to this process is "filling in the gaps" -- that is, extrapolating the expressed intent to cover all the little details that were left unspecified. This system works because humans are at least minimally competent at predicting the preferences of other humans. If your prediction turns out to be wrong, you get feedback and iterate.
Well, guess what. LLMs are also competent at predicting the preferences of humans. LLMs can "fill in the gaps" like no one's business. LLMs can iterate on requirements like no one's business.
Product managers do not speak to programmers in a language that encodes exact requirements, and yet working software somehow gets shipped anyway. LLMs do not need exact requirements either.
I don’t need a model to shit out a REST endpoint. I need it to figure out esoteric errors that take hours or days of debugging. They just don’t do well here. Of course, if a diligent engineer refined considerations from a PM and Engineering Manager I wouldn’t have the job I have.
Now unlike the SRE case, we were the dev team and understood exactly what the logging meant as far as a problem goes, so our prompt started with the correct 0.1% of the system to look at. SRE typically has to start by finding that 0.1% slice from rather more generic metrics. And their interventions have higher risk than a controlled rollout of new code with a specific fix.
But seriously -- newer Claude (and OpenAI and Google and ???) models DO find the smoking gun, if you let them keep going until they reveal the weird chain of events that leads to a bug. I was seeing the most obscure UART driver bug, where it would work at 1,500,000 baud (!) but fail by only outputting the 1st char at 230.4k and 460.8k -- and it was due to a very narrow race that would check the buffer, if not full, insert a character, and return BUT sometimes the TX Complete interrupt would happen between the check and the insert, and something else would insert, and then - buf overflow. At 1,500,000 the other process didn't have time to do that phantom insert. ANYWAY, Claude found this and proposed a fix -- simpler: spins on IRQ-protected buffer empty checks.
I'd hate to think how long it would have taken me to find that.
And THAT's the problem -- of course a human CAN find it, with sufficient focus and time; I'm sure you've found a complicated bug pretty easily sometimes, by sheer luck or good engineering instinct.
BUT, it seems to me, as human, we are capable of creating potential execution paths that EXCEED our ability to EVER figure it out -- due to not being smart enough, not enough time on the problem, or something makes it economically unfeasible.
THIS is where LLMs shine -- let 'em bang at the code for as long as it takes.
The recent Mythos bug-finding explosion I think is proof of this conjecture. I think of it like a chessboard, where a machine really can look at all possible execution paths, and locate obscure bugs; a human programer (akin to a chess program) is doing 'alpha-beta pruning' of what's likely, and only after that list is exhausted are the really weird possibilities examined.
LLMs are our friends. And, as for "WTF did the LLM just do" when it generates code? I always include the instruction "For this code you just wrote, use Best Practices to document this code, function by function and class by class, and when necessary, line-by-line, so that a junior SW developer can completely understand how this code works, using the documentation standard we use (e.g. Doxygen)."
I have also used this technique to learn new languages, or explore ones I only know a little -- it has been a godsend for leveling me up on common lisp, for example. "Give detailed comments explaining what the code is doing, assuming the code reader is fluent in C and Python, and use analogs when possible." Stuff like that.
The reason why a PO can explain something in English, and you get something useful out at the other end (of the developer), is because of a myriad of other decisions you don't see. The reason why some software systems end up being efficient in maintenance and further development, is because of these myriad of other decisions.
The many decisions are the "devil in the details" that LLMs don't get right. Or, let's not anthropomorphize unnecessarily -- LLMs don't know right from wrong, and don't reason or reflect. They could only get this right by sheer luck. In a big numbers game, they'll always get it wrong. If you want to be a PO (or vibe coder, etc) and use English language on one end, and get these details right, there is only one possible approach:
A tight loop with expert knowledge reviewer. The programmer that knows pretty much what they want, in a small section. A LLM can draft it out so that the programmer saves time typing. This isn't really useful for the PO. (PS: The same general advice applies for any other use of LLMs. Tight loop. Expert reviewer)
You'd need a language that can express important details otherwise lost to the English language. And you'd need this to be deterministic. The "myriad of tiny decisions" are the true basis for the code implementation. If they're not expressible in the English language, and they're not achievable by LLMs, there really isn't any other way to achieve them.
I love programming. It's in my blood: my father and grandfather were programmers too. I have written everything from SIMD assembly to Hoon, and implemented several languages of my own. Believe me, I am intimately familiar with the phenomenon you are describing.
It's true, the devil is in the details. And I will grudgingly concede that, at present, humans are better at exorcising demons than AI. But I see no reason to believe that this will remain true. The gap is narrowing rapidly, and even today there are types of demons that AI can dispatch much more quickly and effectively than you or I can. The fact that vibe coding is possible at all (and that people are willing to pay for vibe-coded apps) is proof that an informal English prompt is sufficient to specify software to an acceptable degree. Not acceptable to everyone, naturally, but at least to the creator and the users.
I am not exactly happy about this. It is bittersweet. Much of my identity is bound up in being a programmer. The devil is in the details; but joy and whimsy and great beauty are in the details as well. For a glorious few decades, one could be an artist under the guise of producing economic value. Now, the economic aspect of producing software is being siphoned off, to be done by machines, leaving only the art. I think we will suffer for that, somewhat. But it is a small price to pay.
Yes, feel free to write code. It exists.
In fact, why did you write your comment in English and not code? It's imprecise and doesn't explicitly state exactly what you wanted to communicate, and is instead full of ambiguity and open to interpretation.
Ideally, a cheap verifier checks that the exact requirements are satisfied, rolling back and updating the prompt for another iteration if they aren't. If ten iterations with ten verifications steps at the end of each before the exact requirements are met costs less or in less time than a developer who can accomplish it in one attempt, it is still better.