When you're in a C straightjacket it's hard to take advantages of other architectures like transputer, connection machine, or cell. Ever run bsd unix on a Cray? It was dog slow because the CPU assumed very deep pipelining and always had to continually resynchronize because of the frequent branches in the C code.
GPGPU is the sole recent exception, and even then it's still at the margins. And we have CUDA which is an attempt to, yes, get back to the C model.
These modern CPUs benchmark pretty well, but in practice little code can really take advantage. For example the M1 is a multicore screamer, yet most apps are still restricted to a single one.
CUDA allows different programming models, it's the point of having the PTX virtual machine as the exposed programming model.
CUDA C++ is popular, but is far from the only option people have.
(and it can go very deep, https://www.ibm.com/docs/en/sdk-java-technology/8?topic=egpu... for example)
FORTRAN has tighter restrictions around what you can do in terms of aliasing/etc, which allows compilers to be more aggressive about optimizations without as much undefined-behavior "this looks correct and will run fine until we add another optimization and it doesn't" shenanigans. So in practice it can be faster than C, and it's actually quite popular among the HPC community and other performance-sensitive segments.
It's not like C has to be this way, with all the "undefined behavior" nonsense, it just is an artifact of a programming language written in the early 70s. C makes a lot of sense when a compiler looks a lot more like an assembler with a bit of optimization sprinkled in, and isn't performing super deep introspection and deciding that this entire function can be optimized away because of some arcane "undefined behavior" rule.
I'd say that the most popular option at the end-user level today is a whole level of abstraction above: using Python w/ PyTorch (or less frequently, TensorFlow).
Combined with CuPy and cuNumeric, which is available at https://developer.nvidia.com/cunumeric and handles scaling to multiple machines with Python-written code, through the Legate runtime (https://nv-legate.github.io/legate.core/README.html). This covers a huge amount of what the CUDA userbase uses.
That's why you have things like memory order constraints on atomic operations in C, and why you need to be extremely careful with how you place things like memory fences when doing anything fine-grained between several threads.
Generally, most programmers shouldn't need to be reasoning about OOO or memory order and be using fences. Rather, they should use libraries that are written by experts that use extra-lingual mechanisms (like compare-swap or fences) to guarantee a higher contract.
I mean, you're basically saying that it's only observable if you care to observe it.
I seem to recall that there was a mainframe back in the 1970s that actually did that. Univac 1108 or 1100, maybe? Maybe only in some particular mode?