"OpenAI’s primary bet here has been chain-of-thought monitoring. It is based on an appealingly scalable idea: a lot of the model’s capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives. This does not mean the model will learn to externalize misaligned tendencies that don’t rely on using the chain-of-thought; however, it can allow us to monitor exactly the capability increase from reasoning."
Yes, and it would greatly aid alignment if the user could monitor the chain-of-thought! The open models, including very powerful ones like Kimi K3, are delivering this. The fact that OpenAI and Anthropic are not clearly indicates that a commercial consideration (avoiding distillation) takes precedence over alignment -- regardless of how much they bloviate about the latter.