I'm betting most of the problems are less about variance in execution time and more about measuring the right tasks at the right granularity.
I'm betting most of the problems are less about variance in execution time and more about measuring the right tasks at the right granularity.
The mapper example is actually particularly pertinent. Hadoop's GUI shows me the progress of my M/R jobs with a bar, but it is as dependent on the amount of other tasks running on the cluster as downloading a file is dependent on the traffic in the network. There is no sane way to accurately estimate the amount of time a M/R job is going to take a priori, as far as I know. A stochastic method would be too variable and a method playing clever tricks with psychology seems especially insidious, from an engineer's perspective. You might as well replace it with an "Are we there yet?" button that responds to you in a soothing voice "Not much longer now".