Rather, it’s that when doing the work on one core (with a way simple implementation) you go faster than many of the measurements that popular scalable systems put forward as evidence that they improve on the state of the art. Which sort of invalidates that evidence. Some of the papers end up left with fairly scant evidence, which should be a serious issue for researchers.
GraphChi and X-Stream (http://infoscience.epfl.ch/record/188535) are much more sophisticated implementations, ones you might not be embarrassed to be beaten by, so they don’t make the same point. And they aren’t any faster than the systems above, and so have the same issues (though perhaps less explaining to do).
I'd say, however, the trend is towards data sets being much bigger and I break non-scalable tools frequently in the data profiling process; you can extend the non-scalable ways of doing it by using multiple cores (which is sometimes easy) or SIMD instructions or the GPU. I have even sometimes gone down the rabbit hole of optimizing something non-scalable and hitting the wall. So I am using scalable systems increasingly.
But, I'm also not an expert in graph algorithms, so it's difficult for me to evaluate how much domain knowledge you needed to implement yours.