The 2015 paper is obviously wrong or selling something; it's virtually always IO bound. Yes, I know many people assert otherwise; they're wrong.
Some map reduce loads, especially the kind that people running spark clusters want to do, end up moving a lot of data around. Either because the end user isn't thinking about what they're doing (95% of the time they're some DS dweeb who doesn't know how computers work), or because they need to solve a problem they didn't think of when they laid their data down.
I guess I cite myself, having done this sort of thing any number of times, and helped write a shardable columnar database engine which deals with such problems. If you don't want to cite me; go ask Art Whitney, Stevan Apter or Dennis Shasha, whose ideas I shamelessly steal. FWIIW around that timeframe I beat a 84 thread spark cluster grinding on parquet files with 1 thread in J (by a factor of approximately 10,000 -the spark job ran for days and never completed), basically because I understand that, no matter how many papers get written, data science problems are still IO bound.