Isn't the goal to be able to "debug" and identify alignment issues?
Exactly. This is how the Huggingface incident was reconstructed. But GPT-6 uses partially Neuralese, and its monitorability has dropped sharply according to benchmarks.
Another reason to switch off a central provider to open models as quickly as possible.