I think, as experiments published in last year based on Spark, the inefficiency stems from the assumption that CPU cycles are abundant w.r.t. RAM and bandwidth. So, nobody focuses on optimizing program itself as much as to reduce memory or bandwidth usage.