My personal experience as a software engineer, and a former security researcher who did manual code audit, is that this code review tool has such poor results that it isn't worth the "noise" and friction it causes developers during C/I code review
My personal experience as a software engineer, and a former security researcher who did manual code audit, is that this code review tool has such poor results that it isn't worth the "noise" and friction it causes developers during C/I code review
I found the Devin reviewer to be very good, and have heard good things about Cursor's Bugbot. I've also found asking an agent with fresh context or subagent to adversarially review locally is good.
In a manual review, I then expect all “machinery” to already be properly reviewed, and can focus on design / architecture. I would like an AI assisted review tool to make that part easier, not do the actual review for me.
If I were to automate the back and forth, I would get spurious changes that “fix” what wasn’t broken and a removal of the actual interesting bits.
AI code review is startlingly effective. Continually finding things me and my colleagues never would. Well the decent models do. Maybe not so much the cheap ones.
- review for intent fulfillment: is the ticket done?
- review for correctness: race condition bugs, ...
- review for security: check against this list of sources and best practices
- review for api conformity: identify all surfaces of systems outside this codebase touched by the code changes and check against their docs
- etc. pp., same for maintainability, observability & analytics, test coverage, usage of feature flags
The matrix is sparse, so not every model is used with each of the review categories. Effort levels vary, too. The next stage does a consolidation across all findings, then another stage spins up one agent per finding and investigates the whole codebases for identical / similar instances of the finding; finally, it suggests a fix.
This works extremely well for finding deficits, but the amount of noise drives me insane, too. Lots of feedback is technically correct and "by the book", but pretty useless in practical terms – or even detrimental because the amount of code written and thus the size of the change set explodes. I'm not yet sure how to tackle this problem, any suggestions are welcome!
It is papering over a lower level of competency without having to invest in actual human oversight or real process improvement.
Yes, your prompt need to include to look for certain “quality” aspects you care about. But once that is there it can help find a lot of things.
It can also help in finding edge cases. It is really about the prompt.
I just feel more and more like the effort invested in manual reviews is not worth it
2. If it's the correct solution on a high level
3. Whether it conflicts with or duplicates other parts of the system
4. Whether the comments are actually useful or restating the LLM chat
Also many others but these are the most common IME
The sprawling code comments are becoming the most draining part of code review though, that's really killing me from the inside.
No, none of today's AI would give you enough signal around "should this thing be built in the first place" nor if it's the correct solution on a high level.
They don't understand why you are doing what you are doing, and even if you explain it, they still don't actually understand the motivation and lots of other things.
You'll get them to do guesses and pretend they actually know how to prioritize and will tell you it makes lots of sense, whatever they come up with. But try following it blindly and you'll see where you end up.
This is why "one agent + one good developer" beats "thousands of agents working in a swarm" still today.
> What really important things are human reviews catching in your org?
Another person said:
> 1. Whether the thing should be done in the first place
And you replied:
> All of these are angles an AI reviewer can test for as well, and will (IME) mostly catch mistakes correctly
Which as I noted, is very far from the truth. I neither claimed that you said "agent reviews are a panacea", but when you claim "AI can solve all those things" and two of the first items cannot be addressed by AI (today), then I'm rebuking those specific things, not some other general point you implicitly made.
Where "mostly" means kind of what it says but also not really.
If you don't master this for your own project, what's even the point of your job.
2. Are there potential security, accessibility, performance, etc. issues?
3. Domain specific knowledge (SQL, ASP.NET, XQuery, etc.) where there are better ways of solving a problem, or possible issues not handled.
4. Sense checking ... is the code easy to read? does it need an explanatory comment? does it need named parameters? etc.
Code generated these days with fable and sol are near perfect. What issues they might have is logical errors.
If you're doing a simple CRUD app, sure.
If you're doing anything more involved they get the job done with dozens of shortcuts that bite you in the ass the moment you have on-call duty.
Way too much code and repetition and hacks.
Especially in GPU code, but also in other fields.