I thought we already had reasonably clear evidence that the output in the CoT does not actually indicate what the model "thinking" in any real sense, and it's mostly just appending context that may or may not be used, and may or may not be truthful.
Basically: https://www.anthropic.com/research/reasoning-models-dont-say...