I have been using Claude CLI, Codex, Gemini, and OpenCode (Z.ai GLM-4x/5x) for quite a while. I've seen the code they all produce. I've had to add checks as forcing functions to get them to create more loosely-coupled, readable, maintainable, extensible, testable code. They often get stuck, "complaining" that the rules are too hard to satisfy, and then I show them how, make them take different approaches.
I do not mean to anthropomorphize their work, but I do need to improve what they produce. I think a lot of push-back to vibe-coded projects is that the person asking an AI Agent to make something do not or cannot improve what it produces, which goes back to the question asked. Is it worthwhile to learn programming now. I argue that it is.
So, it would be better if models were trained to produce provably correct code, and agent harnesses are improving to help do this, but the models themselves have been shown to be trained to provide the answer they think you want to hear, and declare a task done as fast as possible. They hallucinate (or lie) about running tests, even about writing code at all (sometimes).
How do cloud models "learn" to write better code? Where are the projects that fine-tune an open/local LLM to "learn" from its mistakes. I have yet to see any model improve on its own.
A lot of the best code I've seen is closed source, and unlikely to be represented in existing training data. A lot of the worse code I've see was online. Maybe newer coding models are getting trained better, but most I've tried are years out of date. Maybe my view is affected by trying so many LLMs. Very few have been great.