I don't work in big data. But here's what I think of your field anyway :-)
In general, you optimize software by starting from the slowest things first, and then working your way to the faster things when you run out of things to optimize. It seems like your problems are I/O bound, and therefore the bulk of optimizations can be made by simply optimizing your I/O (either through async calls, or better understanding of what the frameworks are actually doing, etc. etc.).
And that makes sense for sure.
The thing is: the next level of optimization is not CPU-optimization, but instead memory optimization. In fact, all CPU-based optimization starts at the Main-memory level. Why?
Because memory is slower than virtually everything inside of the main CPU Core.
Which means, memory-level optimization is a far more important skill than any other CPU-based optimization technique. Modern CPU Cores operate at 4GHz easily, but RAM only responds every 100ns on server systems (that's 400 CPU-cycles of RAM Latency!)
-----------
In effect, if there's one low-level optimization any higher-level programmer should know about, it is RAM optimization. Sure, there are CPU-optimizations (SIMD registers, L1 vs L2 cache, and more), but RAM access itself is hundreds of times slower than CPU-speeds and needs to be optimized before lower-level optimizations are sought.
Once RAM access are fully optimized, then you can finally reach for the CPU-core optimizations. Instruction-level parallelism, or SIMD Registers, or what-not.
RAM is the next slowest part of your system after I/O. As such, optimizing RAM access is the most logical next step forward when your programs are running poorly.