Micro-agent: make an AI write code until it passes an unit test
github.com
github.com
i wrote a similar tool the other day: https://github.com/joseferben/makeitpass
it can make all kinds of commands pass by checking stdout/stderr and it’s language agnostic (you need npx to run makeitpass)
For example, have a C function that reads in a file and returns you a string? The string can be checked that malloc actually succeeded, but how do you check that the file actually opened?
Writing a test for that, when it is generally just a call within the function you want to test, isn't really possible. It's not there in the arguments, or the return value, of the function.
How do you check for the right checks in a function expected to do something like this:
int foo() { FILE *fp = fopen("test.in", "r"); if(!fp) { return -1; }
for(int i = 0; i < NUM; i ++){
if(matcher(fp, i)){
fclose(fp);
return i;
}
}
fclose(fp);
return 0;
}I don't really like maintaining tests, it's often a lot of code that needs to be understood and changed carefully
It wasn't, but for starters compilers have always been generally deterministic.
I'm not saying that this is completely useless (I personally think code completion tools such as GitHub CoPilot are fantastic), but it is still early to compare it to a compiler.
I like programming for problem solving, I don't really like writing tests, but that's personal taste, a lot of people like to just use PowerPoint and Jira and tell others what they need to implement, but these people are not software developers.
Whether there's a difference there is in the eye of the beholder, but it does look like that specification languages such as TLA+/PlusCal/Squint or Alloy, or theorem proving languages like Coq (to be renamed Rocq) or Lean look a lot different from the likes of C, JavaScript or even Haskell.
And when things don't work they call people like me, to try to understand the performance problems of something poorly defined and worse written.
Essentially, everyone becomes a red team member trying to think of clever ways they can outwit the AI's code which I for one think this is going to be a lot of fun in the future - though we're still quite a way from there yet!