It just gets “reviewed” by an LLM, which will find a nitpick while ignoring the huge fire in the core of the design, force the planner to make even more sloppy code to cover for an irrelevant test case. Rinse old tokens and repeat until you hit limits.
For example, I recently got brought in to help with quality on a large-scale system that had been ported to a new platform with the help of coding agents. The project was completed and declared operational in record time, but soon after the business discovered that:
1. The promised scalability improvements did not materialize. Instead, it got worse.
2. Observability had been lost. The telemetry was no longer trustworthy.
3. Users stopped trusting it because it was producing incorrect outputs.
What I ended up discovering was that, while it scrupulously kept existing automated tests passing, any behavior that wasn't explicitly covered by a test was free to change any which way. And there were plenty of small things that weren't explicitly covered. Perhaps because the original authors thought they were so obvious and commonsense that they didn't need one, perhaps because mistakes happen. The why doesn't matter. The point is that reality is messy and imperfect, so giving someone a chance to look at things and think, "Huh, that's funny..." is an essential part of defense in depth.
The real worst part was, this whole replatforming was a huge waste of time, anyway. The improvements they were looking for could easily have been accomplished with some controlled incremental changes to the original system. Mostly just removing a few basic and well-known performance antipatterns.
But way back at the outset, the person in charge of the project asked their agent, "What's the best way to X," and the agent gave them a trendslop answer about how Y alternative technology is more scalable and we should just port to that. It was convincing and they were under intense time pressure to just ship some code because leadership is bought into the AI hype and now has the patience of a 4 year old, so they just went with it.
Could be argued they were going downhill before but it's a much faster decline since 2023-ish
Luajit is under 80,000 lines of code.
If I was to employ them to review the code without giving each the same baseline multi-page prompt, they go into endless loop of "improvement" with no end goal in sight.
More and more frequently, I instruct frontier models to stop and go back to the task at hand.
Where does all that supposed productivity go?
The problem is that as a lay user of Firefox, I don't know how you could even make it better in terms of features (I could see it being faster etc.).
https://innovationgraph.github.com/global-metrics/git-pushes
Just because you haven't installed new software doesn't mean that new software doesn't exist.
In fact, having more churn can lead to worse software due to diverging patterns and inconsistency
If I put on my optimist hat for a moment, I hope LLM's would finally make this obvious for people, more stuff !== more value. 250k lines of code a week isn't a flex IMO, it doesn't really mean anything out of context.
So to answer the original question, "Where does all this supposed productivity go?": Sitting in a github repo somewhere, undeployed.
No new browser, no new iOS clone than runs on Android, no new easy to use DaVinci, no new CAD suite, no $5 SolidWorks clone, no redesigned K8s, no 10x performance speedup in Linux kernel.
A couple more step functions in model capability of the type we've seen in the past year, and there will pretty much be no reason for humans to be involved in the development process at all. All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.
It didnt go wrong
And if it did, it was because you werent using the latest model.
And if you were, it was because you didnt have the appropriate guardrails.
And if you did, it's because you didnt have AGENTS.MD.
And if you did, it's because you didnt prompt it properly.
And if you did, it you're still going to be redundant soon because I'm sure the next model released will fix whatever went wrong.
I mean, sure, I could have predicted in what ways an LLM would fuck up, but there's just so many ways I can't keep up.
We just had a major production issue because someone's LLM wrote queries against dev databases. Which are very obviously dev databases because they are labelled with dev in the name, and in the table descriptions. AI reviewer didn't catch it, neither did the human reviewer for that matter.
Kinda what i'm doing already, but for the young startup I'm at that's surprisingly tons of work. I miss the days we wrote code by hand boy those were fun 8.5 hours workdays.
Sounds like the easiest thing in the world: I wonder why did we not think of it earlier?