When you're in a C straightjacket it's hard to take advantages of other architectures like transputer, connection machine, or cell. Ever run bsd unix on a Cray? It was dog slow because the CPU assumed very deep pipelining and always had to continually resynchronize because of the frequent branches in the C code.
GPGPU is the sole recent exception, and even then it's still at the margins. And we have CUDA which is an attempt to, yes, get back to the C model.
These modern CPUs benchmark pretty well, but in practice little code can really take advantage. For example the M1 is a multicore screamer, yet most apps are still restricted to a single one.