Giving LLM’s a <Backspace> Token
arxiv.org
arxiv.org
For it to use the backspace, wouldn't it have to predict the wrong token with greater confidence than the corrected token? I would think this would require more examples of a wrong token + correction than the correct token, which seems a bit odd.
I wonder if this could be have value for when a forward and reverse word predictors collide/fight. I naively assume reverse word predictor models will have the ability to work towards goals, by working from the "solution" back, with the forward word predictor acting as some sort of "causal resolution".
Can't think of a good query to test it on that might need revision half way through...If anyone's got some ideas?
I do find with some coding problems LLMs can start a solution then as it describes its solution it needs to contradict what it has said earlier from it providing/working out itself more context to the problem.
LLM latency is a real problem for interactivity. Especially if you want to talk to it or have it talk to another AI (eg. A video game environment).
When I talk to someone, they don’t wait for my whole sentence to be spoken before parsing and processing and formulating a response. Sometimes they’ll even interrupt me. Can LLMs eventually do this? Can it begin to prepare a smaller collection of “thoughts” as it is fed a streaming input? Or is this at odds with what is actually happening vs. the human brain analogy?
An LLM is "just" a document completer. Given text, it predicts the following text.
So there is nothing stopping you from bolting on a system which works like:
- Given the current streaming input, continually update a pool of N likely continuations (predict what the user will say)
- For each of those N likely continuations, pre-generate responses
- If the actual user continuation matches one of the continuations, use your pre-generated response
This is trading off compute for latency, as you won't use at least N-1 of those pre-generated responses.
An LLM that knows when it should interrupt me, as the parent comment mentions, would be really cool but I don't think it has the ability to determine when an interruption would be helpful. I'd be annoyed if I was asking "what is one plus three plus five" but the LLM interrupts with "1+3=4", for example.
“This reminds me of… that movie where uhhh… the big guy goes to death row and he like… he can help people with his special powers, like that mouse he names or.. uhhh…”
You’d probably interrupt me and it would be appropriate and welcomed.
> This reminds me of
>> A cool spring day?
> that movie where uhhh
>> Star Wars is a movie.
I'm human, so I can understand when I have enough info for a good guess (The Green Mile?) but that's a much different skill than what an LLM does, right?
The bigger issue, in my opinion, is that current models running on current hardware at reasonable costs simply aren't performant enough to do the faster-than-speech prediction that is required to execute the concept well.
They don't do human evaluations of which people would prefer, and instead say that because the MAUVE metric is better with their training method, that the model is better.