The important thing to remember here is there perfect isn't on the table. The benchmark is existing human-introduced bugs vs LLM-introduced bugs. Many developers have encountered odd bugs which a human would not have introduced, while forgetting about all the bugs caught which humans introduced. Or their opinion is formed by models from six months ago.
Try out Opus 5.5 on high. It’s shockingly good. Of course if you’re trying to one-shot a sprawling application with load balanced distributed DBs, you’re going to have a bad time. For small, defined features, it’s pretty fucking great.
We use LLMs extensively on the projects I work in. We don't "vibe code", and we understand every commit.
Fortunately there are people who write software for fun so there will always be some people who would rather do it themselves.
So, I guess people under 18 aren't allowed to learn to program anymore.
They're definitely a useful additional tool for finding more (and more obscure) bugs, but that takes a lot of both human and compute effort too (quite a few of the reported bugs are actually false positives on close inspection, and apparently even with the latest locked down "wonder weapon" models like Mythos), and after all the reports are clean and validated you still can't be 100% sure (but at least a bit more confident) that the code is now free of bugs.
It's not magic and it's not even better or even as good as a mid human, but it's something like infinite man-hours of that drudge work per hour per user.
That will find a lot in old code, and make it a lot easier to keep on finding every little thing right as it's created in new code.