How long until it shows similar results on middle-sized and large codebases? And do the job adequately?
How long until it shows similar results on middle-sized and large codebases? And do the job adequately?
And we should keep in mind that to understand a code change in depth is often just as much work as making the change. When review PRs I don't really know exactly what every change is doing. I certain haven't tested it to be 100% certain I understand fully. I'm just checking the logic looks mostly right and that I don't see anything clearly wrong, and even then I'll often need to ask for clarifications why something was done.
I can't imagine LLMs being used in most large code bases for a while yet. They'd probably need to be 99.9% reliable before we can start trusting them to make changes without verifying every line.