This could allow use some training methods similar to alpha go zero.
I'm thinking if training ai to output edits instead of next token would be enough.
Training would be similar to currently used: just provide some sentences with missing words and asks model to edit those placeholders.
Then hopefully ai would learn how to iteratively built response until it's good enough.
This could allow ai to spot it's errors in the middle of output and start refactoring what it already generated and add some notes (thoughts) that it could later remove as scratchpad.
It could iterate on its context much more so it could have chance to "think" about it