My pet theory is that while the pretraining -> RL pipeline achieves very impressive results, it does not reward clarity of thought or elegance. It's not obvious whether it even should for most tasks, but it does grind on me as a human who needs elegance in order to keep everything under control. You give astra/codex many tasks, it retires them all more efficiently than I could by hand. But you look under the hood and every bugfix is another codepath, it just hammers away at things with admirable persistence and vigor until the tests pass. Similarly in discussions and docs, I've noticed many LLMs like to "beat around the bush."