AutoCodeRover: Autonomous Program Improvement
github.com
github.com
This bug was fixed three years ago in a one-line change.[0] Presumably the fix was already in the training data.
Another thing that seemed odd is the English style used in the responses (watch the video full screen and you can read it).
Ideally, it should also include the problem statement, but that's not in their JSON file and can't arsed to continue working on it – it's just a quick script I cooked up.
I find it very hard to judge the quality of most of these patches because I'm not familiar with these projects.
However, looking at the SWE-bench dataset I don't think it's representative of real-world issues, so "22% of real-world GitHub issues" is not really accurate regardless.
The developer patch for each issue is similarly included as `developer_patch.diff`.
Here are some rules they used to trim down the SWE-bench Lite problems:
* We remove instances with images, external hyperlinks, references to specific commit shas and references to other pull requests or issues.
* We remove instances that have fewer than 40 words in the problem statement.
* We remove instances that edit more than 1 file.
* We remove instances where the gold patch has more than 3 edit hunks (see patch).
And all of this is fine. It's just a benchmark suit and doesn't need to be fully representative. The dataset itself doesn't even claim to be that as far as I can find. All I'm saying that the title wasn't really accurate.
The ArXiv paper mentions the human developer must supply a unit test (which can conceivably be coded with at least the assistance of an AI agent if not autonomously coded, but their experiment relies upon the former kind of unit test) that issues a pass-fail signal. So the 78% of failures are clearly identified, at the cost of implementing TDD for the Issue. The side effects story is punted upon, but I’d still take this over the nothing we have today.
Of course, over a relatively short amount of time using this, I’d expect to experience the 22% (or whatever the real rate is) success rate to drop asymptotically towards zero as the low hanging fruit of the approach are mined out and it becomes kind of like another linter in our CICD pipelines.
The impact of this tooling upon staff skills development will be interesting to say the least.
That being said, when some unit tests are available (either written by developers or with assistance from other tools), AutoCodeRover can make use of them to perform some analysis like Spectrum-based Fault Localization (SBFL). This kind of analysis output can help the agent in pinpointing locations to fix. (Please see Section 6.2 for the analysis on SBFL.)
You have this backwards : it's traditional (at least in the past 15 years or so) to have a test to go along with every code change. The idea is that the test proves a) the bug existed prior to the fix and b) the bug is not there after the fix is applied. Commenters here are noting that ACR generates fixes but not tests.
The patches are tested against a test-suite. So, if there are tests, we welcome them, we definitely use them to validate the patches.
The value I get from copilot is the ability to code faster, not the ability to code.
If tests are available, they can give additional help in setting code context. But tests are not needed, and most of Github issues are solved without tests.
All experimental numbers appear in the arxiv paper. Please let us know if you have more questions.
Strong words!
Not to mention that making the code fix is only a tiny part of resolving an issue. There should also be explanations and added test cases. In other words, I doubt the 22% of “fixes” would pass review by the project owner if a human submitted them.
> The baseline results of Magis (10%), Devin (14%) are evaluated in another subset of SWE-bench, which we cannot directly compare with, so we take the results from their technical reports as a reference.
Wondering how it compares with these models.
/s
edit: apparently it's not the exact same benchmark but a similar one
https://github.com/nus-apr/auto-code-rover
if you need example bugs we can provide that too. Some examples also appear in the arxiv paper, please see
I think I speak for others when I say the best way to judge the efficacy of this project is some real-world, on-site examples of it being used in prod. I'm especially curious for its performance in feature-request or flakey bug report type issues as opposed to reliable test failures. I expect the former is much tougher!
- From sympy: https://github.com/sympy/sympy/issues/13643. AutoCodeRover's patch for it: https://github.com/nus-apr/auto-code-rover/blob/main/results...
- Another one from scikit-learn: https://github.com/scikit-learn/scikit-learn/issues/13070. AutoCodeRover's patch (https://github.com/nus-apr/auto-code-rover/blob/main/results...) modified a few lines below (compared to the developer patch) and wrote a different comment.
There are more examples in the results directory (https://github.com/nus-apr/auto-code-rover/tree/main/results).
Which excludes 80+% of real world bug and feature issues, in my experience…
The auto-generated patches are to reduce the effort of resolving issues. In practice, they should be reviewed and verified by human developers before they are integrated.
The local code search idea to get around context limits is cool. Have you experimented with Anthropic's models for the larger context limit and dropping the code search?
https://github.com/nus-apr/auto-code-rover
Please try it out and send emails to the contact emails in this webpage, if you have any questions.
Is it expected to be able to solve arbitrary (simple) bugs, or only the list of bugs in the benchmark set?
If someone gets 20% on an exam I don't go "great, thats 20% of the way there!!!", instead I go "you clearly didnt attend, try again next time".
Sure, if little Bobby gets a 20% I’ll whoop his ass, but if the inanimate hunk of metal on my desk gets a 20% I might start to take notice.