LLM use language, but it can't "think" about biochemistry
I saw that LLM have reasoning capabilities, which is different from machine learning, but I don't understand how it works.
LLM use language, but it can't "think" about biochemistry
I saw that LLM have reasoning capabilities, which is different from machine learning, but I don't understand how it works.
There's also something to consider with lower level vs higher level abstractions in language. E.g. jargon. One short word could have a 200 page thesis behind it defining all the ramifications. Talk about compression of information.
Now imagine if our language lacked say the mechanism of jargon, of using some meta word to define thousands of stringed together words at once. Every idea like "car" would have to be described from first principles. The species would probably never develop technology with this sort of language pattern present. If we could somehow level up beyond our current abstraction level, maybe that would make us even smarter, able to handle bigger ideas quicker in real time.
Even more simply than all this: I can only speak about what I have english words for.
It’s a combination of cultural assumptions, facial expressions and affectations, thinking patterns, and a whole cultural upbringing that can lead you to very different mental processes and natural conclusions starting from the same words and phrases.
Language absolutely encodes a certain form of intelligence. A lot of those things are reflected not just in the totality of the culture but the language itself. Being fluent leads you to different thinking patterns and different conclusions when processing in that language.
Now that we are once again attempting to unify our language we find ourselves in a pursuit to build something to escape the Earth.
https://inv.nadeko.net/watch?v=Or_3tlEOLj4&pp=ugUEEgJlbg%3D%...
“In the beginning was the Word, and the Word was with God, and the Word was God.” - John 1:1
https://en.wikipedia.org/wiki/Logos
There's a lot more to it than just "language."
Anecdotally, but I have lived in different cultures with entirely different languages and/or dialects, and the thoughts and even entire categories of thoughts people from these cultures express, or can easily express, are very much shaped by their language. Relatedly, I've also often witnessed multilingual people switch out of their native language to a second one just to express a particular idea or nuance, because they can do it with two words in that other language but would need at least a couple of sentences to say the same thing in their native one.
We use formal language to express symbolic relationships, e.g. "A implies B". But even "A implies B" has multiple meanings: material conditional, strict implication, logical entailment, etc. So, symbolic systems are not "pure and hard", they are also contaminated and softened by the vagaries of language outside them, which is our primary access to those systems: "valid" natural language and its strings of words. A statistical system that can string words into valid(=allowed by the distribution) language asymptotically approaches reason. So, the mind is not in the words, but in the laws that permit many words to come together, i.e. the probability distribution.
Modern chain-of-thought models with RL post training on verifiable tasks + realistic environments + rubrics are worlds apart from models trained on a simple next token prediction objective.
More money goes into the rubrics and RL environments than individual training runs themselves.
(Yes, at inference-time LLMs still output words one at a time, much like human speakers. But don't confuse the mechanism with the training objective.)
/s
Incorrect. As the OP said, that is a very 2023 understanding of how LLMs work.
Grab a new model from OpenRouter. Have it work on a task. Change a few tokens and have it continue the completion.
With or without a harness?
Have you actually tried this yourself? Of course it can derail it. Try to reflect on your interactions with LLMs without all the constraints like web search, agentic scaffolding, etc.
The same way that a “yes” or a “no” input from you can change the response, cot tokens are fed back into the model as input and can derail it.
It would be interesting.
I have seen such derailments within the GHCP harness maybe with GPT 5.6 Luna that went into some loop about whether it already provided a final response to the user, or 5.6 Sol suddenly switching to talking about MS SQL performance.
I also saw a post about Sonnet unexpectedly talking about Minecraft after seeing a file with a related name. The user thought it was the output of another user's conversation so the post was fairly popular.
Thank you for making my point for me. But let’s keep the goalposts stationary. We’re talking about LLMs without scaffolding.
I still don't know if that is the case, and how frequently it happens, since you did not share details beyond vaguely suggesting it would happen.
Does this mean that a single incorrect word or twitch will completely derail the task you’re trying to performance? Or will you, like any other intelligent being, recognize it and compensate?
With reasoning models, a derailed chain of thought can be rerailed.
What rerails it?
This realization is something you assign meaning to. For the model there’s no difference between either of these states.
https://www.earth.com/news/our-brains-are-constantly-working...
https://www.psycholinguistics.com/gerry_altmann/research/pap...
https://www.tandfonline.com/doi/pdf/10.1080/23273798.2020.18...
https://onlinelibrary.wiley.com/doi/10.1111/j.1551-6709.2009...
And I also found the video I was referring to https://www.youtube.com/watch?v=FHQfmJEpRmU
I have a feeling we know more than that about how it works.
We didn't build our brain.
Typically when you build something you have a decent idea how it works.
And that means we are not privy to whatever things it has learnt in its trillions of weights.
We might have built them, we sure as hell didn't design them. And no, we do not have a decent idea about how it works.
~"Predicting the next token is not an insult. It's pretty much what we all do."
I don't know about you, but I tend to speak one word at a time...
That isn't how tokens work, nor is it a representation of how the brain represents information.
Please predict the next word.
Intelligence is implicit in language understanding. The best possible next-word-predictor is omniscient.
That wasn't too hard, maybe I'm superintelligent?
Omniscient for the set of "meaning" embedded into it's training set. It's not broadly omniscient, big difference.
Just for kicks, I actually put your sentence into an LLM. The response was along the lines of, "Your query was incomplete and about medical knowledge, so I need to be careful. There is currently no cure..." and then goes on to do a decent job of summarizing existing treatment approaches for metastatic breast cancer.
What's so interesting about this is your notion of prediction here is divining the answer in reality, i.e. finding a cure for breast cancer. But its notion of prediction is determining the next logical sequence of words given its training set, so it produced a block of useful and context-relevant text, but not what you actually care about. This leads into the much broader question of what do we mean by "intelligence," which forms do these things have and not have, etc. etc. If nothing else it's all very fun to think about and debate.
Train on a massive body of text, figure out what correlates with what, and next thing you know you have a rather impressive facade of logic that can even connect things in novel ways where a connection is clearly called for, but not yet made. I call it a facade because LLMs will be able to advance knowledge significantly in finding these clear connections, but they exist only because no human can hold more than a tiny percent of all knowledge in their own mind.
Where I expect they will run into issues is in finding the unclear connections - like going from an existence where math doesn't exist, to one where somebody 'invented', or more aptly - discovered, math. That's inventing something from nothing, rather than just logically connecting pieces. I don't see how this is possible with a token prediction algorithm.
Anyhow, the point I'm making is that language itself includes encoded logic. And so LLMs working as token prediction algorithms are able to exploit this functionality to produce statements that offer a facsimile of logical reasoning under a constrained domain.
What is it that makes something truly novel or creates something from nothing?
When we do it, do we apply existing concepts, combine them with a general intuition for how physics work in the real world, and use that to form a hypothesis that we then test in experiments?
So try putting yourself in this ancient mindset before mathematics. How did somebody invent it, come up with the concept of numbering everything, further develop the various 'tricks' for manipulating these numbers, and so on? In terms of raw 'complexity' it's far less impressive than the latest LLM models solving some obscure mathematics problem that almost nobody understands.
But in terms 'intelligence', I find it vastly more impressive - because it's again this sort of difficult to describe concept of going from nothing to something. There is no logical baseline that naturally and cleanly leads to math. Almost like a child would say when asked how they learned something, 'Oh I just thought it up.' Except in this case, somebody genuinely did!
Provided the training data was extensive enough and training rewarded solving problems that require mathematics.
I also don't think the people behind the LLM companies think this is the case either. If it were then it'd make so much more sense to drop the current regime and instead move to the most basic systems trained on nothing but the most fundamental first principles and have them try to derive everything from there. It'd ostensibly lead to far more reliable systems with little to nothing in the way of bias. It'd also likely be vastly cheaper than the current practice of trying to train on essentially all consumable knowledge.
sometimes with residual connections, but we can ignore that for sake of simplicity.
Promoting LLMs is encoding the problem we want into the query vectors, and through the magic of the complex training and the power of operations in a very large dimensional abstract space the AI can manipulate the representations, and iteratively approximate solutions. (And using bigger and bigger contexts and better encodings it can form better models.)
Not sure how it is now, but early “reasoning” was simply the big labs sticking “wait a minute, what if I…” type language blocks into the process to trigger something like our own internal reasoning.
Machine learning trains the network to do... anything that you reward it for. If you keep training, it keeps getting better.
Next word prediction can always keep getting better.
At first, simply "learning" spelling is what makes the predictions better because tokens are word chunks, not always whole words.
Then, the models "run out of steam" and can't get any better by learning more spelling rules, but the gradient descent forces them to get better... so they do... by learning the rules of grammar.
At this point the AIs can output correctly spelled and grammatically coherent sentences, but the sentences ramble on about nonsense topics.
So what happens next as the models run out of grammar rules is that they're forced to learn the rules "above grammar": logic, world knowledge, coherent story telling, etc.
At some point they learn to output pages and pages of fluid, coherent text, but... if they're not smart, if they don't think, and if they don't know what they're talking about, then they're still "suboptimal" and their forced gradient descent will make them close those gaps.
Eventually, the only way they can improve at "next token prediction" is by building up to human-like intelligence, including an inner monologue, theory of mind, and everything.
We can even read their "thoughts": https://transformer-circuits.pub/2026/workspace/index.html
Some human was looking for something like this once. They didn't find it, but they wrote about the search precisely enough that the finding can happen during inferrence.
Maybe somebody will come along and school me, but for now it's a fun way to think about it: A million dead ends, each with a uniquely disappointed human, now with a chance at a second life in the hands of a different human they haven't met. If only the weights had encoded enough to introduce us, supposing they still live.
It’s weird, I’d be hesitant to say “it’s a voice” but it kind of is and it is not my own which I find curious (who on earth is speaking in my head). In some ways it sad, if I close my eyes, I can’t picture a sunset and I can’t really dream. I love reading books but I can’t visualise the settings properly but it resonates with how my mind describes the world to itself.
I don't. I seem to think at a more abstract, pre-verbal level rather than through an internal voice.
Some studies suggest that frequent internal monologue may occur in roughly 30–50% of people [1], but the research is based on relatively small samples.
[1] https://www.psychologytoday.com/us/blog/intersections/202304...
I think something like this proved to be true when it came to folk theories of different learning styles (e.g. visual vs language based) once those started being tested in rigouros ways. People could still be right but I would be interested to see what we get if we test more directly for subvocalization or fmris for language based brain activity.
Also saved pesos on the charge-per-text SMS schemes the local phone companies used because we could embed information across so many options.
"The words of the language, as they are written or spoken, do not seem to play any role in my mechanism of thought. The psychical entities which seem to serve as elements in thought are certain signs and more or less clear images which can be "voluntarily" reproduced and combined....From a psychological viewpoint this combinatory play seems to be the essential feature in productive thought....The...elements are, in my case, of visual and some of muscular type. Conventional words or other signs have to be sought for laboriously only in a secondary stage, when the mentioned associative play is sufficiently established and can be reproduced at will."
Nope. You up and pee.
Most human reasoning happens within language - even mathematics is an abstraction that allows us to map concepts we don’t natively hold into a linguistic processing layer.
We can, at best, approach a good set of weights, even in tiny neural networks.
Imagine if we found a way to calculate the exact optimal weights for a given loss function. I mean, there is an exact optimal solution, it exists, but we can't find it exactly, even for a neural network with just 50 parameters.
There is an optimal set of weights that minimizes the loss function for a given set of training data, but we cannot find it.
Granted, even if we could, it might just be overfitting.
So I'm not sure how it knows to be 'surprised' that alone is pretty fascinating.
It’s sort of like all the people who will ask Claude or GPT to validate their complete nonsense and receive unyielding praise for it, the models just learned that this is the best received response based on training data and RL.
I bet these same sorts of expressions can be found in practically every failed attempt as well.
I mean, the subtlety of the neural network weights that emerge from training are not fully comprehended by anyone, man or machine.
Every individual calculation is understood, and every step of training is understood, but the exact nature of those weights that divide the responsibility of responding to subtle changes of input in intelligent ways is beyond me.