On a superficial level, if you are doing overnight processing of log files, then you probably care more about throughput than latency. In this case, averages are probably a fine metric. On a slightly deeper level, standard deviation is only a useful measure if the distribution is known, and in a lot of real world cases it is not. The right question isn't whether 100 or 1000 tests on the same data provides sufficient statistical power, but whether range of inputs is sufficient to trigger worst case perfomance.
Now, I presume that Zed knows these things and applies them appropriately, but the article strikes me as more snide than helpful. Perhaps as others say he's a great guy in person, but I prefer my stats with less attitude and more insight. Here, for example: http://yudkowsky.net/rational/bayes
[edit: changed my sloppy language from 'has no meaning unless to the distribution is normal' to 'is only a useful measure if the distribution is known']