Finding Compiler Bugs by Removing Dead Code
blog.regehr.org
blog.regehr.org
The PDF it links is 404, but the Wayback Machine caught it in a few different revisions (the date of Regehr's post would correspond more closely with the earlier revision):
http://web.archive.org/web/20130810232543/http://pdos.csail....
http://web.archive.org/web/20131214080638/http://pdos.csail....
Another idea: taking bug reports for one compiler for a language and using that as fuzzing input for another compiler for the same language.
The C language needs a good test suite with known expected results and good coverage of the language's definition, but one way to build this suite is not to catenate all the regression tests of existing compilers. Actually, part of the reasons why we do not have such a reference testsuite is that it is not as easy as putting together regression tests from various origins.
2- ???
3- get around to making a conformity assessment test suite for C.
We forced Xuejun Yang (who turned Randprog, the prototype that came before Csmith into Csmith) to fix more bugs than were necessary for his PhD or than you can expect a research prototype to have bugs fixed (I am one of the developers of Frama-C, pitting Frama-C “against” Csmith was my hobby for a summer, and we found and reported as many bugs in Csmith as we found bugs in Frama-C). The sentence “over the last couple of years I’ve slacked off on reporting compiler bugs” near the post's conclusion is telling. You can expect the same story to unroll for EMI. Researchers are not rewarded for maintaining software ad vitam æternam, even if the software is useful, but hopefully they have or will soon release the generator as open-source, and then if you find it too useful not to use further, you can fix or work around bugs as you find them.
As an example, I think that there is still one bug in Csmith that we work around by ignoring programs that have the symptoms that usually indicate that bug (we still use Csmith to test Frama-C after we have finished a major feature that could introduce the sort of bug it can detect).
See also http://blog.regehr.org/archives/1058 on the same blog for additional thoughts on the fate of academic software.
This type of research spawns other research and projects outside of academia by acting as a proof of concept, even if the original researchers stop reporting bugs.
For instance Csmith (and others) inspired Gosmith, which has found a number of bugs in the Go compiler. I hope that someone will use the obviously successful strategy of EMI to improve it further.
https://code.google.com/p/gosmith/ https://code.google.com/p/go/issues/list?q=label:GoSmith
What I left implicit is that a non-reduced Csmith test makes a terrible regression test. It may for instance spend several seconds incrementing a counter from 0 to 4000000000 before switching to the entirely unrelated computation that once triggered a bug. The value of these randomly generated tests is in generating a new one the next time, not in saving them to run again and again. Running the same randomly generated tests again and again would find very few bugs and would be a criminal waste of electricity.
The reduced program is worth keeping as a regression test, because it is typically a few lines long and these few lines contain a construct that once tripped the compiler. Sometimes the reduced version can be rewritten by a human to be even more concise and readable than the output of C-reduce. But as I said in another comment, one compiler's regression test does not obviously make a good test for another compiler.
Property-based testing is, as far as I can tell, a meaningless term since all testing is property-based.
http://blog.jessitron.com/2013/04/property-based-testing-wha...