R1's technical report (https://github.com/deepseek-ai/DeepSeek-R1/blob/main/DeepSee...) says the prompt used for training is "<think> reasoning process here </think> <answer> answer here </answer>. User: prompt. Assistant:" This prompt format strongly suggests that the text between <think> is made the "reasoning" and the text between <answer> is made the "answer" in the web app and API (https://api-docs.deepseek.com/guides/reasoning_model). I see no reason why deepseek should not do it this way, if not considering post-generation filtering.
Plus, if you read table 3 of the R1 technical report, which contains an example of R1's chain of thought, its style (going back to re-evaluating the problem) resembles what I actually got in the COT in the web app.
The result after that could actually look different though for usual questions (i.e. summarised in a way chatgpt answers on questions would look like). But it is usually very coherent with the code part, so if for example it has to choose from two libraries - it will use the one from the reasoning part, of course.
If by the "visible reasoning" is just for show they meant these models don't actually think and reason, then yes that is correct.
But if they meant that the visible reasoning is not quite literally a part of inference process...that's entirely incorrect.
R1 is open source. We don't have to make guesses about its functioning.
I have seen it infer incredibly obscure things in the chain of thought that I was impressed it could piece together.
It is an incredible tool. I would trust it 1000% more than a random person on reddit.