Training language models with pause tokens
arxiv.org
arxiv.org
From the future work section:
> better determining the number of <pause> tokens (perhaps using model confidence)
If the number of pause are learned, using model confidence + regularisation (so that the model doesn't always use the maximum number of pauses), then we effectively have the System 1/2 switch. If it's a task the model has seen tons of times before, it just goes with the first inference pass. If it's low confidence, then it keeps expending more inference passes until it reaches an acceptable confidence threshold.
> Readers of “Thinking: Fast and Slow” should read the book as a subjective account by an eminent psychologists, rather than an objective summary of scientific evidence
https://replicationindex.com/2020/12/30/a-meta-scientific-pe...
While most of the psychology field is crumbling and filled with bullshit (even though most people didn't catch on yet). There's still _some_ truth lying in there.
It might look bad in isolation, but Kahneman's work is one of the few that actually holds up to scrutiny.
This post by Scott Alexander is pretty good in summarising the few good bits left. https://www.astralcodexten.com/p/heres-why-automaticity-is-r...
---
But regarding the initial point. The goal is not to fully emulate human thinking. It might sound wishy-washy, but there's _obviously_ _some_ truth to the system 1/2 model. It's not perfect! But I think it's useful.
We as humans do most decisions without thinking (hard), but we have a way to _switch_ into a more reliable but expensive mode. We see some evidence that artificially inducing LLMs to 'think' more improves their output.
So it stands to reason that adding this capability to a model would make it better. And what better way to do it than swallowing the bitter pill and have that decision be made by the model itself, on a case-by-case basis by learning it from data.
Yes, not all people think like that. But that also means that there are people who do think like that, based on the application of the pigeonhole principle on your own statement.
All avenues for internal computation should be done to improve model accuracy & results, including this method.
Klein’s model consists of two steps:
1. Recognize the course of action that makes the most sense
2. Imagine how this would look in reality"
I, for example, may produce something like the following sequence of internally represented words and non-words "A voice narrates... [imagines what it can feel like] Does it feel external to [non-verbal reference to conversation partner]? It doesn't work like that [implied reference to myself] [pause for introspection] I produce words. It's not like someone else narrates thoughts in my head. [stop producing words to check that] I just think, and one thing leads to another and that's why it's hard to stop the sequence of the words being produced by me and it's not because I can't stop some intrusive voice inside my head" and so on and so forth.
Phonetic realization of the words isn't important if I don't specifically concentrate on it.
I think this is the same misunderstanding as “the mind’s eye” https://x.com/stuartjritchie/status/1708871999924179024?s=46
From the tweet I linked earlier:
> This is just an issue of these things being extremely hard to explain using language, and people using words in subtly different ways.
In the paper, the model doesn't really learn to use padding tokens effectively till after training so it's the only thing.
For instance, I tested whether asking a model to produce "a b c d e" before solving a math problem improves performance.
is of course not what you would expect to work. It is not about dumping meaningless stuff into the input but intermediate results. If you have to add two long numbers but are only capable to add single digits, you want to dump the carry into the output and get this fed back as input so that you can take it into account when adding up the next two digits. Dumping unrelated stuff into the output will of course not help. Chain of thought does exactly this, it solves part of the overall problem and feeds that back for consideration when tackling the next piece of the problem.
Certain tokens will guide it to more accurate responses just as much as others derail it out of distribution. This should be obvious.
Literally this paper, you can see the model only really learns how to use the padding tokens effectively after training.
What does not work?
CoT is so obviously not just about padding for extra compute.
In a certain way it is but the crucial point is that an autoregressive neural network is pretty much a recurrent neural network and chain of thought encourages the use of the feedback path.
I'm saying that even if you do what you suggest, it simply doesn't work that well. If you read the paper, you'd see extra computation from padding is not used effectively till after training.
Yes, extra compute helps with CoT but it's not even the reason it works. I believe CoT works moreso because it nudges the computation in the right(er) direction.
This starts with 0 0 6 1. What do you have to do to generate the next number? Count the numbers in the output to figure out that you have done 4 numbers, add 1 because the 5th prime after 53 is next, figure out that this prime is 73, calculate the digit sum 10 and then find the remainder 3. There is just no way that a neural network has learned to do this.
But ask for this as a table with prime, digit sum and digit sum mod 7 and it becomes trivial. First prime after 53? Output 59. Digit sum? Output 14. Mod 7? Output 0. Next prime after 59? Output 61. Digit sum? Output 7. Mod 7? Output 0. Next prime after 61? Output 67. Digit sum? Output...
This is what chain of thought does, it allows the neural network to only do simple tasks it knows how to do, output the result and continue from there with the next step. You would keep those intermediate results in your short-term memory but a large language model has none and so the best thing it can do is to dump them into the output and read them back when generating the next token.
The proposed pause token has essentially the same goal, the neural network can dump some state into the output to work on it but this output is not considered part of the response. It is essentially an attempt to hide the chain of thought process and only respond with the final answer. The hard part will of course be to train the neural network to make efficient use of this, to actually output useful intermediate results before the pause is over and the response gets extracted.
1. Train without pauses for fast response
2. Train with variable numbers of inserted pauses, so the model can take more time (computation) to improve its performance.
Where pauses are represented both by pause tokens, and a new pause countdown element.
I.e:
Sequence:
…, token, pause, pause, pause, token, …
Pause element: …, 0, 3, 2, 1, 0, …
3. Train the model model to generate its own next pause value (insert its own pauses), to optimize a meta-performance tradeoff function that ranges from the regular performance measure (independent of delay) vs. delay.The tradeoff is based on an “urgency” value between 0 and 1, that adjusts the tradeoff meta-performance weighting between just delay-independent accuracy vs. just speed.
The urgency is used at each step to calculate the tradeoff meta-performance, and as a new urgency input element so the model knows what tradeoff level it is supposed to optimize.
4. Train the model to generate its own urgency values, based on examples of prompts which either communicate urgency based on request or context.
“Quick, give me an estimated value for …”
“The reactor is about to explode, what are the shutdown codes?”
Note, that the model generated urgency value and the next urgency input don’t have to be the same.The generated urgency could be pushed higher or lower algorithmically, to get the next urgency input, if the model is being used in a context that has its own urgency relevant information.
—-
So many good ideas. It’s clear models are going to improve quickly.
Next step beyond a model that can leverage pauses for better performance, and a model that controls its own pauses, is a model that can be interrupted.
Interruptions could be urgency related, cancellations, modifications, or irrelevant:
“Hurry up”
“Forget that, do this other thing.”
“Oh, and the answer should be odd.”
“While you think, I am enjoying a nice cup of tea.”What’s worse, memory use is quadratic in sequence length so adding 10x tokens is just not very smart.
You don't have to use 10. That's just what they found optimal for a particular benchmark. For others it was lower.
>What’s worse, memory use is quadratic in sequence length so adding 10x tokens is just not very smart.
Everyone uses flash attention these days. Memory is no longer quadratic in scaling.
The paper should have answered this question and well. It is an obvious question.
WRT FlashAttention, surely you understand that the matrix-matrix multiplication must still happen even if it’s a fused kernel that doesn’t need to hold several intermediary variables in memory.
>WRT FlashAttention, surely you understand that the matrix-matrix multiplication must still happen
Of course it happens. The point is that it doesn't scales linearly now with respect to memory.
I don't understand. You have an S by S matrix. You increase S, the matrix grows quadratically. There is no way around this.