Thinking mode is exactly this. The LLM is allowed to dump thoughts, correcting itself along the way, and only surfaces the cleaned-up final answer to the user
It's not exactly this. It's something like it. But it's not a token that erases a previous token. Not that I know how that would work or how you'd get training data (edit histories of internet comments seem too few).
(You could have a token that hides previous tokens, but that'd be rather closer to CoT.)
Backtracking, of sorts. The LLM descends a seemingly fruitful path, and all of a sudden that path seems less so. So you backtrack a few tokens ("erasing"/"undoing") and resample the probability distribution a bit back in the stream.
Would it retain some memory of the fact that it backtracked? That path becoming less likely somehow?
Yeah that's what I was imagining. No idea how though, so this is just rambling on my part :)