> Thinking traces should be treated as black boxes. There is no point in reading them.
Just a few minutes ago i was reading Qwen 3.8 27B's reasoning when i asked it to do something that was computationally intensive and it started going down the rabbit hole of doing it using some GPU acceleration approach - even after leaving it to "think" for a bit, it never realized there is another and simpler way. So i stopped the generation and added a "note" saying that as the problem is computationally intensive, it could become much faster if using an alternative approach.
At least in my experience (with local LLMs, i don't know how the cloud stuff behaves) what LLMs "do" tend to correlate with what they "think", so being able to read what they "think" is valuable.