How to add a million bugs to a program, and why you might want to
moyix.blogspot.com
moyix.blogspot.com
We attempted in the paper to make an argument based, effectively, on whether the modified program looks like a normal program from the point of view of a bug-finding tool. From that perspective, the control flow in the program is unmodified, and the data flow is exactly the same as in the original program except for a small portion where data from the input file flows through a global variable to the crash site. We show that this section of data flow is usually short (in terms of number of instructions in a dynamic trace); we're also working now on extending it so that the data flow is more "natural" (i.e., it goes through return values and function parameters, rather than through global data). Additionally, we showed that the bugs are distributed throughout the execution lifetime of the programs we tested.
I definitely agree that they do not look like bugs humans most often write. However, to be effective, we just need them to look enough like the kind of bugs we want to catch that bug-finding software will have to reason about the program in the same way.
There are a number of types of bugs that are significantly more common than others, so it's possible to draw equivalence classes around them and use them as representatives of the whole. That's the basic technical skill around testing in general so isn't terribly scary as part of a SQA process.
It runs the test suite, then introduce bugs and test and at the end compiles a line-by-line summary of what bugs where injected, caught and missed.
It actually sounds reasonable to me. Of course it won't necessarily produce the same bugs as a human programmer, but it's yet another automated quality tool, in this case, a sort of "test of your tests".
In all of testing, proving a negative is futile. But if this tool can create bugs that your bug-finder cannot find, and your real dev process is at all correlated to the bugs this tool creates, then this tool is a useful addition to the toolbox.
Alternatively, this randomly mutates programs, which means no time was lost thinking of what kinds of bugs to introduce. Then, I cannot see how the bugs introduced are representative of real-world ones. It just, as another commenter said, would be fuzzing your tests. Useful, but it doesn't tell you anything about the number of remaining bugs in your code.
Thanks for saving my time.
Note: Especially for severe ones that can lead to code injection.
So far some of the things you'd expect hold true: fuzzing doesn't tend to reach bugs deep in the program, symbolic execution has trouble with path explosion and one-way functions, etc. You can see those results in the tables in section VI.D of the paper.
My thought was: These scripts are running with set -euxo pipefail. I can insert a 'false' line between any two lines, run the script, then remove it and run it again and I should still get the correct results. I never finished writing this test tool, but I still think it's a good idea I want to get back to.
Nope, not related.
haven't used it for anything realistic, so not sure how well it pans out in practise.