The author seems to trivialize how hard it is to write a compiler that generates efficient code. Like designing a simpler language than C (less general architecture targets) would solve much.
It's not about designing a simpler language than C, but rather one that matches what hardware is actually doing better. The example in the article is how GPU programming works. To make the most out of modern hardware we need to change the way we program to embrace concurrency, parallelism, and eschew shared global state entirely. The way FP works is much closer to hardware than the imperative style.
I have a hard time figuring out how forbidding concurrent task from communicating would necessarily make them faster. Also, how is FP a more accurate model of eg. AMD64 than imperative style? How would eg. "Clear Global Interrupt Flag" instruction be modeled in FP with no states?
I don't know if it is a more accurate model of AMD64/X86 but it may be a better match for the underlying silicon (gates, logic circuits). I can see function pipelines and composition of these in some ways analogous to circuit design. After all those instruction sets in AMD64 are merely an implementation of these. Pure functions and pipelines could model circuits quite well.
Shared state is not a problem with current hardware, mutable state is also bit a problem. The only issue is with shared state that is mutated from multiple CPUs and that's why the single writer principle is a fundamental rule of multithreaded programming.