I think https://arxiv.org/abs/2412.06769 is a good description of the premise of reasoning in latent space, although https://arxiv.org/abs/2604.15726 argues it's already what really happens.
It's as if the only way you could think was by writing down a word, erasing all the thoughts from your head, then reading the word you just wrote down and deciding on the next word, etc.
Reasoning purely in latent space means that the model would still produce an output equivalent to tokens but unconstrained e.g. the output could be raw and opaque vectors. A significant downside is that you lose the ability to inspect the reasoning trace. It would also make the reasoning trace potentially larger which has operational issues.
Not completely true: KV is a projection of the activation at each layer's input, so attention heads see (a representation of) all previous tokens' activations at that layer. The hard decision at the LM head doesn't change that.
What if instead of a latent they used at a shorter form of note-taking-like reasoning, using more symbol to achieve a denser CoT. We'd get the best of both worlds. WDYT ?
You can see evidence of this phenomenon in models dating back to the OG Deepseek R1. It was common to see the model talk itself out of the correct solution in the <thinking> block, or fail to reach it at all, only to produce a correct answer in the response. And vice versa; it was also common to see it reason its way to the right answer and then fail to follow through in the response.
I can see the reasoning being a substrate for computation, but in which space should we interpret this computation to be happening? The vector representations of individual tokens are completely different (and even the way the reasoning traces are broken up into tokens will be pretty different) between Qwen and Claude e.g.. The only way I can see this being effective (which it is) is thus that we SHOULD interpret the model to be "computing" in natural language and thus we can indeed take the chain-of-though somewhat literally.
The Deepseek R1 behaviour you describe is from a model from january last year, are you sure this is not pathological behaviour rather than an indication of the reasoning not needing to be taken literally?
I do however agree with the point that it is not necessarily a dead-end. That Qwen loops almost at an OCD-like level, but retains accuracy on the times it does answer, shows that. Yes ideally it loops less, but I am for now happy to accept that this is what it takes to run models locally. At least it is available for our inspection.
Fine tuning or post-training is effectively biasing certain outcomes: making them more likely to occur. This comes with trade-offs. A coding LLM will bias technical language, which would harm a model for general use.
This opens a really interesting field of research. Our brains use specialised regions because specialisation turned out to be the most energy efficient method for biological compute. It might also be the best performant. We don't want to activate 100% of our prefrontal cortex to breath. What a stupendous waste of the organ. I think we see incredible advancements in model clusters in the future, using specialised models for specialised tasks. We have the appearance of this today in some harnesses, but they are shallow imitations. The real innovation will be low-cost, accurate routing. Existing solutions are woefully inadequate for many reasons.
Now though I'm considering all the hidden "thinking" in the models layers that happens for each token output. It is a wild amount of waste! We just can't see it.
This kind of stupid excessive computation is fundamentally how these models are so good.
One day hopefully not so soon someone smart or a foundation model will come up with a more efficient architecture. That's when things get really scary.
The important part of "actually wait, I really need to XYZ" is just "XYZ".
The model can attend to just "do XYZ" and produce almost the same vector modifications as full verbose "reasoning".
It's good to remember that LLMs have no more state then what they can derive from the context up til any point. So if that context is hard to interpret, that will reduce effectiveness.
LLMs are trained on human natural language, not a specialised internal-only monologue to make syntactic shortcuts. Their response should make grammatical sense to a human reader because they are mimicking human speech.
"foo bar" is ambiguous.
"pursue theory foo; no, this didn't lead anywhere, let's backtrack and pursue theory bar" is accurate, should be meaningful to LLM attention, but is too verbose.
The minimal caveman way to say this is "not foo. instead do bar".
This really should provide the LLM attention with everything it needs to grasp the intention, but with far fewer tokens used.
Instead of all this I could have just said:
"nothing ambiguous. clear but less words"
And this memory control primitive leaks into the reasoning chain, because it has no other channel for it available and we do not know how to train any other channel.
On the flip side, it tends to converge quickly, roughly proportional to the actual difficulty / clarity of the task.
I am very confident the reason we get all these second guessing and "but wait" and "actually" is they train them on collapsed corrected sessions. i.e they take sessions that look like this:
user: Do x.
agent: the user wants me to do x. I think I need to do a and b first.
agent: does a.
agent: does b.
user: No no no doing a was wrong you should do c before b.
agent: undoes a. does c.
agent: does x
And they turn it to a session where the user correction shows up in the thinking. i.e user: do x.
agent: the user wants me to do x. I think I need to do a and b first.
agent: but wait maybe I should do c instead of a
agent: does c
agent: does b
agent: does x