I disagree. They do reason during the reenforcement learning stage. They don't reason at inference. A good metaphor is that useful output are like nuggets that exist after reenforcement learning which need to mined to be, in LLM talk, "surfaced." Without supervised fine tuning, the reasoning models will add weight to tokens, words and phrases like "verify" and "check work" which will cause it to follow those verifying tokens with reasoning tokens that do just that, verify.