I wonder if giving the models context of the temperature of its past generations would help here. Like a thinking mode that deliberately has a section that is high temperature, while the rest is lower.
If we could make progress in that area, maybe CoT could gradually decrease as it approaches its limit, or maybe the LLM could control the temperature of the next token itself (how this would be trained, I have no idea).