Seems like they are closer to scratch than reasoning... Generating some scratch to draw from helps make it easier to compute the real answer.
What is less apparent is that humans do.
This seems indeed one obvious hole in the argument of the paper. There is no indication whatsoever that human thinking process is more reliable than LLMs intermediate tokens. Which doesn't make our thinking useless, as messy as it might be. We reorder and explain it after the fact.
They are trained by gradient descent, but inference doesnt involve it.