In classification task, MSRA and ReCeption show very similar performance: 0.03567 vs 0.03581 top(5) error. The gap is much more drastic on the localization task: 0.090178 vs 0.195792
The residual learning presented by Microsoft Research seems to be a breakthrough, based on the early evidence. But yes, ImageNet needs to be updated to stay relevant.
FWIW, "human performance" is not as meaningful as most people make it out to be, because this is such a narrow task.
They used a dataset from 2012. Are you saying that the test labels from a 2012 dataset have been kept secret until now? Really?
Baidu was banned. I don't think Facebook has ever done ImageNet?
But I agree that most of the interesting work will happen outside ImageNet now that human performance has been comprehensibly surpassed.
The only exception is non-Deep Learning based systems, where some people remain convinced that alternative approaches can match DL systems and have other advantages.