I'm currently working on porting a mid-sized project to a new architecture, new programming language and of course adding new features.
Getting a new feature implemented is quite easy. You spend a few hours brainstorming specs with the agent, then ask it to implement it. This gives you extremely frequent code drops that add a new brick, add a new feature, etc. All of this with 100% code coverage (we also have mutation testing, strongly-typed code, standard and custom linters, etc.)
Then you look at the code. Code that has passed review, generally. You realize that the database schema has been broken silently, and that the agent has rewritten the tests or the golden fixtures to match. You realize that it has made assumptions that contradict the specifications and the product is going to break once it's in the hand of users. You realize that the 100% code coverage is essentially a convenient lie, because the code and tests have been written to make passing easy. You realize that none of the security golden rules have been followed, and that has managed to happen because the agent has somehow deactivated linting.
Why did it pass reviews? Well, because of deadlines. And because there is simply so much code (and so much unparsable/misleading documentation) that it's simply impossible to review all of this. And because things move so fast that nobody understands the CI pipeline anymore, and the explanations of the agent are convincing enough that surely, it knows better than you?
On the upside, bugfixing becomes so fast! Just add a new test, wait a few dozen minutes, and a new Merge Request appears. With equally convincing/misleading explanations, and something else broken.
After ~4 months, we had a bare bones deliverable, which we're now steadily expanding. If we had had to write the product manually, I suspect that it would have taken us at least one year, possibly two. So, that's the productivity increase. The productivity decrease is that what we have is not a product but a glorified demo, something that will work very nicely on the happy path, but on any other path, all bets are off.
> "Why did it pass reviews? Well, because of deadlines. And because there is simply so much code (and so much unparsable/misleading documentation) that it's simply impossible to review all of this. And because things move so fast that nobody understands the CI pipeline anymore, and the explanations of the agent are convincing enough that surely, it knows better than you?"
I have come to realize that AI is so "successful" because the system in which it is being deployed was designed to push product as fast and cheaply as possible from the start.
Humans are usually overworked and stretched to their breaking point, which I originally saw as the source of our broken software woes, which, like our streets in the US, just get a new layer of asphalt to cover up the crumbling bits each year instead of rebuilding the infrastructure with reliability and longevity in mind.
My former employer was using both Claude and Codex for firmware that was driving an over-burdened power circuit that itself was partially designed with ChatGPT. All of the individuals involved approach LLMs with god-fearing reverance because they do not understand _how_ the LLM works, just that it _does_ in a "good enough" way and they can offload their thinking, which is something we all wish we could do because thinking is hard, time-consuming and costly. I get it.
But like you mentioned, tests were being passed, not because the code was sound, but because the tests were altered to match the results. This is not necessarily the fault of the agent, either; it's just interpretting the prompt(s) - written by a flawed human, btw - with stochastic mechinations that seem to make a great deal of sense on the surface, but remain unable to be followed or repeated by the brains of (most of) its users.
As a rresult, I had to deal with product that work great in the field...at least at first, before it start literally catching fire, ruining its own powertrain because everything the agents touched became too complex with too many subtle cracks in the veneer to review properly. The system (read; capitalism) demanded viable product quickly to please investors, and the burnt-out humans who decided to try this AI thing ended up trusting it nearly completely, so any ideas of repeatable and complete testing, diagnostics and root cause failure analysis morphed into a sloppy "it works on the bench" checklist before being sold to a customer who had come to trust that their deceptively simple product would just work as advertised.
I'm going to die on the hill that AI as a replacement for our brains is precisely how we will make ourselves go extict, but I am old enough to already be regarded as a crufty dinosaur who is stuck in his ways, and I'm made peace with all of that. What I can't get my head around is watching people use this awesome tool (and it is, admittedly, awesome) to literally just speed up all the mistakes they were already making. Perhaps it is because I am aging, but slowing down and having a think seems more valuable to me now than it ever has, especially when creating something new. AI is powerful and, like any good tool, could be useful in the right hands, but more often than not I see it being used as an accelerant for all the worst parts of product development to appease a market that has suddenly been told they can now pick all three points on the Iron Triangle instead of just two. This makes about as much sense to me as taking a laxitive when you already are suffering diarrhea.
> and the explanations of the agent are convincing enough that surely, it knows better than you.
I feel this in my bones. I also get to watch the misalignment feedback loop close itself when the next agent sees that security rules aren't followed because of a hallucinated 20 line justification in a code comment, and then it decides that the project _is_ a demo and then confidently writes even more security holes into the codebase.
Then when you catch the issue, the agent pushes back against the fix because it would need a schema change and production DB migration.
They are very good at that.
> Then you look at the code. Code that has passed review, generally. You realize that the database schema has been broken silently, and that the agent has rewritten the tests or the golden fixtures to match.
I find that the code review leg of this is critical to invest a lot of energy into hardening.First is don't trust the code review from your local harness, even if it uses sub-agents; externalize it into another system.
Second is to get your most critical human code reviewers to encode their heuristics in markdown files and feed those to the code review agents.
Third, if possible, is to bring the code review "into the loop" so that it's not only running in the PR, but also running in the coding loop so the coding agent has immediate, external feedback. Final PR code review is a backstop.
This pattern [0] works well because it solves for some team level problems where folks are using different harnesses or different models (consistency issues) and it means that code reviews don't just sit at the end of the loop; it actively alters the code production cycle.
At the end of the day, LLMs are still tools.
Convenient way of saying "you're holding it wrong".
But that's boring and you can't build a YouTube audience around it.
LLMs are just tools indeed.
I don’t know what to say to those types anymore. Live and let live I guess.. or in this case, not live I suppose.
Useful talk and meetings all day is fine, but these are just to write busy/billable hours for all these useless folk with no value for the project. And socially I like listening and talking, just not for this.