Many researchers would love to do just that, of course. However, as many researchers will lament, it's not always easy to get 10+-figure node and 11+-figure edge data sets appropriate to a space being explored.
The best research benchmarks I see do compare to a single-core (often multithread or multiprocess) implementation. And they also show benchmark results on datasets of increasing sizes.
I agree with the authors that those sorts of papers aren't common enough, though. And we should strive to do better. Moreover, I agree that in practice in industry, many people over-optimize for horizontal scalability early, and/or do not realize potential savings and benefit by doing vertical optimizations after gaining initial scale.