If you had multiple people look at your PRs multiple times on different days results would be very similar.
If you had multiple people look at your PRs multiple times on different days results would be very similar.
It’s not perfect but usually it works pretty well, and I’ve had the model come back to me with oh actually the test passed, the bug doesn’t work exist
As a bonus, you’ve now got a test that can detect that bug if it comes up again.
The "keep improving" the code base prompt have been tried and it never works. The LLM has no consciousness of where to stop and where to draw the lines of reasonableness.
typically this means there is some ambiguity in the specification, and the model flips between alternative interpretations
For a normal review loops you can ask the model to return with nothing found if nothing is found and not invent things and it will do a better job of exiting without anything found.