My instinct says that these systems will expand their complexity to fully fit the cognitive budget of the agents that coded them and then atrophy the same way human-built systems do at lower cognitive budget. Only this time, because of the larger up front budget, the complexity ceiling will be higher, and the potential depth of the problem may be much much larger. It may mostly manifest as increasing cost over time - the agents grind for longer and longer, iterating over and over to fix all the failing tests, and the breaking point will be where it never converges and you come back to millions of dollars in budget spent and still tests are failing and effective gridlock on system changes.
But this may be all my human-biased fantasy that justifies still taking a role in software development.
wow this is a beautiful way to put it
This thing called attention economics has captured my attention quite hard. Digital opiates are everywhere since 10 to 20 years. And in my view, the part that they implicitly of explicitly strive to capture your attention is new.
To use the wooley term “quality”, the top 20% might stand a good chance of making huge strides. But the remaining 80% of projects (in particular the bottom 20%) will atrophy extremely quickly. Yet, these will be the project that many push LLM’s too as their domain/technology is complex and/or outdated. Digital transformations that can be done quickly will be tantalising but ultimately unsatisfactory long term (as you describe).
How many are you seeing / estimating?
Everyone is fatigued by endless code review which you get no credit for and has become massively more of a burden.
All PRs are superficially fine now. There are no typos, there is unit test coverage, but there are deeper issues that require massive amounts of effort and time to spot.
Lots of review comments about various conditions that wouldn't feasibly happen (same shit with claude now).
But then I'd see these same reviewers approving PRs where the bigger design was just fundamentally broken. Oh, we're adding a blocking call on our hot path, but at least the method name makes it very clear that it is blocking.
In general I agree that the current AI reviews are creating too much noise and it is masking these bigger design issues.
I'm old enough to remember using CVS and then subversion in companies. People would commit straight to main (which was then called "trunk"), because making feature branches and merging them was cumbersome. And, on regular intervals, the person responsible for some corner of the codebase would do a show-and-tell presenting it to peers, but without the sharply defined boundaries of what the code looked like before vs. after some recent set of changes. People might remember some things from the previous show and tell or from first hand experience with that code, but that kind of memory is necessarily fuzzy, and diffs weren't an artefact that was typical to look at. So, these reviews didn't block people, and any comments that came from reviews defined a direction that things should go from here on out. If a corner of the codebase was deemed to be in a bad shape, the blame around that was equally fuzzy.
https://github.com/josephmisiti/awesome-machine-learning
It's helped a lot. Agents haven't figured out how to do that yet, or sendgrid, sns, etc are doing the hard work for me.
In the meantime, for business communication, I use AI to shorten my text, to make it more concise.
So far I haven't had a reason to go back through commits to isolate any issues but if I do hoping the 'why' messages may come in handy for my LLM lol
Until I fully understand what's going on, the PR doesn't move and my interrogation of the LLM doesn't end. My interaction is littered with "Explain X" and "How does this square with Y?" and "What if Z happens?"
The interrogation is the point, without me having to wade through hundreds of lines of irrelevant code to get at the meat of the matter.
I can read the room. Coding is going to go the way of some other disciplines where machines do most of the detail work and and we know stuff works by verification. There are other fields like this.
Do I like it or not? That doesn't really matter. I need a job, so I'm going to get good at the new way to ensure I continue to have a job. I consider this to be a smart thing to do for myself and my family.
Unfortunately I see a lot of (senior as well) engineers who think that just a vanilla LLM reviewing another LLM is sufficient, and my comment was directed towards those. If however you see the LLM era as needing more test support and systems than ever before in the form E2E tests and so forth, where "code review" as such becomes mostly irrelevant as you have such a strong test system in place that if that passes you can be sure it doesn't break anything for users, then yes, that's good.
Code reviews may not even happen, or if they do, it'll be all automated, and the verification will be the key.
Ask anyone in the semiconductor industry when last they understood the design of those things.
It's very easy to argue with stawman arguments.
We have this at work : fully AI-generated code and description. People will give review comments generated by AI which the "author" replies with an AI-generated response, all with LLM wording full of jargons no one understands not even the person who sent it. When you ask them what they meant, yeah idk Claude said so
Every time the LLM throws jargon around, you call it. "What do you mean by gated wedge?" You call its bullshit, check what it's saying against your understanding of the overall system, and keep it on the straight and narrow.
It's a lot like supervising a junior dev who happens to be very quick at absorbing lots of info, but not so great at the big picture.