In some ways the reality is worse: the same version of an LLM won't get better at its job, even though you might get better at prompting it. Newer versions are trained on more code, which has obvious benefits, and their harnesses are better at taking advantage of tools that were created to keep human coders out of trouble.
The other side of the coin is that coding agents are not maximally productive unless you give them enough rope to potentially hang themselves. Over roughly the past year, the coding agents I use have gone from hot garbage to pretty consistently useful, especially if I find tasks where I can give them a lot of running room. On the other hand, last week I found a case where the coding agent was looping and flailing like it was doing every third try a year ago.
They fail less often, but they fail in the same way.