Training LLMs to Reason in a Continuous Latent Space
arxiv.org
arxiv.org
I am not surprised at all that Meta was able to generate some positive returns by feeding the last hidden layer back into the model auto-regressively.
The method of training they describe in the paper is really cool. Summarized in Figure 2, they train it with a corpus of step-by-step text instructions and then across multiple stages, they iteratively replace one of the textual steps with a last-hidden-layer embedding and see what the model spits out. The weights are then updated through cross-entropy loss as the additional text tokens are generated once again.
So they're basically rewinding the output, replacing an increasing number of textual steps with hidden state embeddings, and playing it forward as the model gradually learns to do all of its step-by-step thinking using just the hidden state data.
In a way, this might be how humans learn to think through language. Our parents teach us using words and our brain gradually replaces the words with thoughts until we can replicate the action or solve the problem ourselves without anyone guiding us with words.
In this manner, the model makes use of the rich semantic information encoded at the last hidden layer while informing the user via an extraction of that hidden layer specifically tuned to generate human-legible concepts such as, "I'm considering the impact of converting the units from kilograms to pounds," or whatever.
I've said it before, but I think it's just something like Quiet-STaR, but simplified. They have a bunch of question answer pairs, many of which are difficult. They generate a lot of tokens from the question (let's say, 3x the length of the expected answer), summarise whatever is generated and reinforce whenever it generates the right answer.
I don't think o1 is something complicated.
> the last hidden layer is obviously super rich in semantic information
I don't agree that this is obvious, and think it's likely wrong (see the sibling thread [1]). The model has to at some point compress down its prediction for the entire future string of text to a prediction for a single token. There's no prior reason to assume it does this mostly in the final "LM head" linear layer, and the inputs to it don't have to predict anything other than the very next token so there's no reason it should (which is what I think psb217 was getting at), but I'm not familiar with what research has been done into it. On the other hand, processing seems to typically be concentrated in the central layers.
That vector has a lot of information in it, it's not a debatable thing.
As noted above in parens, look at the llama 3.x models. The space is already shared in some sense. It's called "tied embedding".
Encoding the next token is the minimum possible amount of information it might contain; that's not much information (the distribution over the next token is just a projection from the embedding space). E.g. it would be useless for any classification task.
In the case of llama 3 its 4096 * 16 bit = 8192 bytes of information...that's like 8192 characters of ascii. More than enough for most classification tasks... and if you jsut spend any time thinking about encoding the logits for a vocab of 128k... you'll come to the conclusion it's likely to require at least several hundred bytes (maybe 1000?) to do it in any way that will actually work in practice.
Many existing applications make use of hidden layers in a transformer to perform useful tasks such as classification. The concept of an “embedding” is simply the output of a hidden layer, after all.
I found this out a few years ago and I was shocked that the way I read wasn’t universal. I have since been practicing reducing/eliminating subvocalization and I am getting better, and it allows me to increase my reading speed significantly. It also serves as an excellent example of how different our internal mental processes can be, and how completely unaware we are that there could be any other way to think than our own.
Reading speed slams to a speech-paced crawl until the reader can "not think of a pink elephant" again.
iirc this is a well supported task iirc called "classification head" instead of "language modeling head" in case anyone else wants to do this as a fine-tuning project
This paper makes me wonder whether, in a very fuzzy sense, we could give #LLMs access to some similarly crystallised analog of emotion or emotional valence, below the level of language
OpenAi as a mega-organism might be intelligent, but the LLMs definitely are not.
The "compressed capture of semantic relationships" is a new thing we don't have a word for.
"Tokens" don't have to be words, or woofs...
[1] It still doesn’t make much sense to view this as a statistical process; dogs can generalize far better than transformers, as perhaps best seen with seeing-eye dogs. I believe dogs’ powers of causal reasoning exceed what is possible from mere surface statistics: e.g. they innately understand object permanence as puppies, whereas transformers still don’t understand it after viewing thousands of dogs’ lifetimes of experience.
But then going the other way, from statistics to a causal model, is just an inverse problem -- just like, say, going from a set of noisy magnetic field measurements at the boundary of a container to a pattern of electric current flow inside a volume, or going from planet positions to orbit shapes and periods to an inverse square law of gravity. Generating a compressed inverse model from surface statistics is exactly the sort of thing that deep learning has proven to be very good at. And by now we've seen no shortage of evidence that LLMs and other deep networks contain stateful world models, which is exactly what you'd expect, because for all their parameters, they aren't nearly big enough to contain an infinitesimal fraction of the statistics they were trained on.
So I think it's overly dismissive to regard LLMs as mere surface statistics.
It's literally what they are though.
Yes those probabilities embed human knowledge but that doesn't mean that the LLM itself is intelligent. It's why every LLM today fails at anything that isn't centred around rote learning.
You can predict better and better with simply more knowledge i.e. data.
Seemed easy, and I thought harmless, tho maybe not
> We use a pre-trained GPT-2 (Radford et al., 2019) as the base model for all experiments. The learning rate is set to 1 × 10−4 while the effective batch size is 128. Following Deng et al. (2024), we also reset the optimizer when the training stages switch.
Basically, the lm_head (a MxN matrix where M is the dictionary size and N is the internal dimension) can be seen as the dictionary too. You can think that and the softmax over it as compute cosine similarity of the last hidden output w.r.t. input embedding dictionary.
In that sense, they are sharing the representation space.
(BTW, I believe sharing lm_head with input embedding not working as good as separating them, so only mobile focused LLMs do so. So here is that. It would be interesting to experiment if injecting a projection layer like you suggested would improve performance or just red-herring).
Do they? Interpretability techniques like the Logit Lens [1] wouldn't work if this were the case. That author found that at least for GPT-2, the network almost immediately transforms its hidden state into a "logitable" form: you can unproject the hidden state of any layer to see how that layer incrementally refines the next token prediction.
[1]: https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreti...
Like in a CNN for instance, if you fed later representations back in to the first kernels they wouldn't be able to find anything meaningful because it's not the image anymore, it's some latent representation of the image that the early kernels aren't trained on.
I feel back pain when I read the crazy, unsound speculation about how the brain is supposed to be like a computer. Serious mistake.
I think we lack a good formal definition of what (fuzzy) reasoning is. Without it, we will always have some kind of unexplained hallucinations.
I also believe AGI could be implemented as a model that can train models for specific tasks completely autonomously. But that would kill the cash cow, so OpenAI etc. are not interested in developing it.
The term "computer" has it's original outside of "electronic computer". It used to be a role, a job function. There has been no time in human history where the only computers have been electronic computers.
But, sure, let's be more precise: Any Turing complete system is equivalent to any Turing complete computer and can reasonably be called a computer, but let's also limit it to any system that can not compute functions outside the Turing computable set. We don't know of any such systems that have been shown to compute functions outside the Turing computable set, at all, including brains.
The rational basis for assuming the brain is a computer is that we have not a single shred of evidence that exceeding Turing computability is possible, nor any theory for how to even express a function that is computable for humans but not Turing computable.
If you can find one single such example, there'd be a rational basis for saying the brain isn't a computer. As it stands now, assuming it isn't, is nothing more than blind faith. .
if your basis that anything that has equal to or less than turing computability is a computer, then everything is a computer.
It could still turn out to be intractable without a better architecture, but the notion that it might not be impossible makes a lot of people very upset, and the only way it can be impossible even for just an LLM with a loop bolted on is if human brains can compute functions outside the Turing computable set.
> Any Turing complete system can emulate any other Turing complete system, and an LLM can trivially be made to execute a Turing machine if you put a loop around it
Wouldn't it be more efficient to erase the LLM and use underlying hardware as Turing complete system?
BTW. Turing test is just admission that we have now way of defining human level intelligence apart from "you'll know it when you see it".
Also we may have to look for better loss functions than ones that help us predict the next token to train the models if the objective is reasoning.
Could someone explain the last hidden state of the LLM ? What it shape is and how it is normally used - and why it hasn't been used yet to augment the next input? (which seems logical)
There's typically an "unembedding layer"/"classification head" that uses this hidden state to produce a softmax distribution over the LLM's vocabulary. In this case, we can think of this as "snapping" the hidden state into a single token and feeding that token into the next position of the autoregressive LLM.
In this sense, the last hidden state _does_ augment the next input. The authors simply propose directly feeding this hidden state into the next step rather than reducing it into a single token—thus, reasoning in continuous latent space rather than discrete token space.
Don't forget, non-linearity is fundamental to the whole process, otherwise you'd just have one large linear transformation. Maybe there's a similar role for discretization? :shrug:
tldr; the method requires training in a way that loses one of the major benefits of transformers, but maybe in some scenarios that loss is worth it.
Essentially, English is one of many possible encodings of an underlying intuitive, possibly non-symbolic representation.
That's debatable. Language shapes thoughts much more than you might think. Because you learn concepts from language that you could not imagine by yourself until you learned/read about them, so they are in effect very linked to language.
For example, I can think in formal logic. I've learned to do that, and surely my brain takes a step-by-step approach to it, but I've also internalized some of it and I don't think that my proficiency with English has anything to do with it.
I could have learned the same concepts in any other language, but the end result would be the same.
And surely there are many thoughts that can't be expressed purely with words. For example all that is related to qualia. You can think of a color but you can't describe what you see in your mind's eye with words, not in a way that would let a blind person share the same experience. Or try describing "love" without making a similitude. Is love a thought? Or a feeling? Is there a meaningful difference between the two?
This embedded concept is what is doing the work, the word was just the seed of the understanding and a method by which to convey that understanding to others.
In all cases though these thoughts are kind of tied to representations from the real world. Sort of like other languages via different senses. So yeah, how abstract can our thoughts actually be?
https://en.wikipedia.org/wiki/Linguistic_relativity
>The hypothesis is in dispute, with many different variations throughout its history.[2] The strong hypothesis of linguistic relativity, now referred to as linguistic determinism, is that language determines thought and that linguistic categories limit and restrict cognitive categories. This was a claim by some earlier linguists pre-World War II;[3] since then it has fallen out of acceptance by contemporary linguists.
(universal grammar peoole hates this somehow, it's weird)
If you don’t mean for safety reasons, I’m not sure why.
I'm not sure whether it wouldn't be more reliable to let the model run on latents and try to train a separate latent-reading explainer module that has at least some approximation of what we want as an explicit optimization objective.
Assuming it actually is or has the potential to be better than CoT, from what I gathered from the paper the current results are mostly just more efficient token-wise.
Predicting the next word is not intelligence.
It really seems like we are building a true intelligence, adding components to different parts of a "brain" until we have something rivalling the human mind. It's exceptionally dangerous and it's remarkable how researchers turn a blind eye to any possible consequences.