Humans are not born being able to achieve that. It is learned behavior. You and everyone else will remember their teacher saying, "Check your work!" Both the human child and the model work through multiplication problems using the same technique, using the distributive property. They both try to get a reward. For the human child, it is a sense of someone commending them for correctly solving the problem, a reward that probably yields some type of positive dopamine or serotonin feedback loop.
The model solving the problem will have a lower error rate if the first series of tokens created is followed by a series of validation tokens that are subsequently followed by error-correction tokens if there is an error!!!
Maybe it is thinking. Maybe it is remembering to validate and check the work and then remembering to fix the error. For the model trained with RL, why did tokens associated with validation towards the middle of a stream of tokens yield much better results? DeepSeek proved with R1-Zero that a model will learn to verify and correct itself from RL alone with no supervised fine tuning (SFT) teacher ever showing it how. The only reason DeepSeek used SFT was to clean up the reasoning tokens to be human readable. [0] When constrained by SFT, the models will use the double meaning of words -- polysemy -- to satisfy being human-readable while also carrying meaning for what they are working on.
Different people think differently. I watched a viral video of some ~11-year-old child talking to his mom or dad about a stream of a voice in his head. He discovered for the first time that he has a stream of thought. When he goes to school and solves a long multiplication problem, like the stream of tokens from the model, that voice will say to itself (him), "Check your work!"
That is a case of the stream of thought as words being aware of the stream of thoughts as words. Self awareness is a different conversation.
What I think is happening is that the child's stream of thought while solving a multiplication problem in school is likely very similar to an AI model's stream of tokens solving a multiplication problem. And they both were learned. The mechanics are very different, yet, the analogy is apt.
[0] https://huggingface.co/chutesai/DeepSeek-R1-NextN/blob/main/...