What's the difference between using GPT to write the prompt to GPT, and "thinking"? The LLM uses the first tokens to predict more tokens, and then uses those tokens to predict even more tokens.
Part of it is as a another comment in this chain mentions the chance to review the prompt. Part of it is that it forces the AI system to plan things in a certain order, in much the same way that forcing the contractor to write the plan out first forces the contractor to proceed in a certain predefined order that may (or may not!) be better at getting to a final answer.