And frontier models routinely crush all the above in a way I couldn't, at speeds unattainable to mere flesh and blood like me.
People making end-user applications might think they can tolerate more errors and bloat from AI.
Just because they can get away with doing that with AI (and that's debatable) does not mean that people can also get away with that in developing tools, libraries, languages etc. The errors, bloat and instability bubbles up exponentially as people build on it.
There seems to be this fallacy of "I don't have to write code anymore, therefore nobody will have to write code anymore."
I see a lot more of "AI coding doesn't work well in this specific case, therefore it's entirely useless".
Library-type work has mostly been side/toy projects, although fwiw, with a standard/spec on hand (CommonMark for example), I'm also happy w/ the output. It's often possible to "close the loop" and have the coding agent autonomously iterate until the standard is adhered to.
Creating something that is solid enough for widespread, reliable building is just in another category. And I wish people recognized this distinction more when they say we don't need to look at code anymore.
That said - I would (again, maybe naively) suppose it's not hugely different - much of the work I do occurs in code where many people have and will work on it, and where the size of the codebase dwarfs model context windows.
In that case, I feel the same - current frontier models, when properly oriented to a task, with some assist on the big-picture thinking - are more than capable of generating good code that can slot into big codebases with many moving pieces. Of course, I'd have to point to other people's work to defend this, but I think that's still pretty reasonable especially against the declared "LLMs are worse than useless for generating code".