They are not able to comprehend that for anything more complicated than that, the code might compile, but the logical errors and failure to implement the specs start piling up.
Grok 4 Fast told me its own internal system prompt has rules against autonomous operation, so that might have something to do with it. I am having decent results with it though.
I've noticed Claude shutting down conversations on the same subject. It says "you may continue this conversation with an older model." ChatGPT also got extremely uncomfortable talking about it, and refuses to build anything in that direction, which I find amusing since its ancestor GPT-4 built a self-modifying Python programmer in 2023.
(Also, OpenClaw meets GPT-5's definition for "extremely dangerous AI software" due to the self-modification factor. But I think that applies trivially to any agent, so...)