One thing that I'd add (somewhat selfishly as it relates to my PhD work), is the idea of generating datasets that are deliberately challenging for different algorithms. Scale this across a test suite of algorithms, and their relative strengths and weaknesses become clearer. The caveat here is that it requires having a set of measures that quantify different types of problem difficulty, which depending on the task/domain can range from well-defined to near-impossible.