I've been developing a suspicion that many people simply aren't looking at what their agents produce and can't speak to the quality of it in a fully informed way.
For context, I've set up harnesses with recursive automated review loops, spec driven development, explicit lists, better models, better harnesses, formal methods tooling, etc. All of it helps, but they don't eliminate output issues. Those become very apparent when I go through the slow, manual work of deeply comprehending / validating LLM code.
And that leads me to one of three conclusions. Either my standards are achievable only by hyperintelligent programming gods, I have a skill issue using LLMs, or others aren't applying the same level of attention.
The first is obviously untrue. I meet my own standards and I'm an idiot. The second seems unlikely because I can see my competent coworkers and well-regarded people in the community discussing the same issues. So that leaves the third.