The big problem is that AI output can be very convincing and look "right", even appear to work, until you examine it in detail and realise all the edge-cases it didn't handle.
The big problem is that AI output can be very convincing and look "right", even appear to work, until you examine it in detail and realise all the edge-cases it didn't handle.
My own reviews have shifted now. I tend not to examine detailed semantics anymore. The AI is as good or better than me at assessing whether a chunk of code does what the author said it was supposed to do.
Instead my job is to spot design and architecture smells, broader semantic errors, violations of unspoken business requirements, etc.
For example, I was recently reviewing code that built out a transactional flow. Part of that flow involved recording the transaction somewhere user visible and I knew that should only happen after the transaction was confirmed. AI implemented it where the transaction was posted.
That starts as an issue of underspecified requirements but that always happens in the real world. Thus that's where I can provide the most valuable insight: assessing with that context, whether business, operational, historical, or forward looking.
In my experience, I find it to be exceptionally good at exploring and fixing the edge cases. Of course the output is not human maintainable for these fixes and needs to be heavily tests controlled, refactored or just accepted as being agent-maintained going forward.
We still have a very strong review culture, and people work hard to review their own code before making PRs to avoid wasting other people's time.
Its a massive effort for our team it has to be a massive effort around the globe if you take quality series.
It will just be easier better cheaper to teach ONE LLM how to do good code review
So…are you planning to hire only seniors who already know everything? Where do these seniors come from?
The "engineers" using that AI. Could they repass the whiteboard tests of old? Have not some of them quietly left?
> What remains to be seen is whether Google also introduced more Chrome bugs in June than over the past two years, thanks to human programmers.
> The big problem is that human programmers output can be very convincing and look "right", even appear to work, until you examine it in detail and realise all the edge-cases it didn't handle.
Your argument falls apart quickly because the exact reverse is equally as likely. Also, whenever I see someone say "just wait longer" for some effect to become apparent, it feels like a lazy argument.