I like the idea of using fuzzing "to show a direct correlation between programs that score low in their algorithmic code analysis and ones shown by fuzzing to have actual flaws".
I hope they'll use AFL, and publish the parameters / settings, so others will be able to repeat the experiments.